3 papers
cs.CL2025
Essential-Web v1.0: 24T tokens of organized web data
Essential AI, :, Andrew Hojel +22
Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible…
cs.LG2025
Practical Efficiency of Muon for Pretraining
Essential AI, :, Ishaan Shah +22
We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon…
cs.CL2024
The Impact of Visual Information in Chinese Characters: Evaluating Large Models' Ability to Recognize and Utilize Radicals
Xiaofeng Wu, Karl Stratos, Wei Xu
The glyphic writing system of Chinese incorporates information-rich visual features in each character, such as radicals that provide hints about meaning or pronunciation. However,…