No content.

Written by
More to read
Chunked and Fused Cross-Entropy: How Online Logit Tiling Slashes Large-Vocabulary VRAM Bottlenecks in LLM Training
Chunked and Fused Cross-Entropy: How Online Logit Tiling Slashes Large-Vocabulary VRAM Bottlenecks in LLM Training As frontier large language models have scaled, tokenizer vocabularies have expanded substantially. Where early architectures such as LLaMA and Mistral relied on 32,000 subword tokens, contemporary models routinely employ vocabularies of 128,256 tokens (Llama 3), 152,064 tokens (Qwen 2.5), and 256,000 tokens (Gemma 2). Larger vocabularies compress text more densely, improve multilin
1 minKakao Splits Into KakaoAI and KakaoX to Accelerate AI and Messenger Integration
South Korean platform giant Kakao Corp. announced a corporate split that will separate its core operations into two independent publicly traded entities: KakaoAI and KakaoX. The restructuring, approved by Kakao's board of directors, aims to isolate and accelerate the company's artificial intelligence engineering and messaging ecosystem from its broader investment portfolio. Under the spin-off terms, existing shareholders will receive shares based on a net asset book value split ratio of 36% for
1 minUS Warns 35 Partner Countries to Choose Between Pax Silica and China's WAICO AI Coalition
The U.S. Department of State is preparing formal diplomatic notices instructing 35 partner nations to select between Washington's AI alliance and Beijing's competing framework. According to a draft cable reviewed by Reuters and reported by The Decoder and CNBC, the U.S. warns that countries joining China's newly established AI initiative will be excluded from the U.S.-led Pax Silica coalition. The diplomatic draft states: "To be part of everything is to be part of nothing. Signature of the Pax
1 min


