4 papers
Essential-Web v1.0: 24T tokens of organized web data
Essential AI, :, Andrew Hojel +22
Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible…
Practical Efficiency of Muon for Pretraining
Essential AI, :, Ishaan Shah +22
We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon…
The Impact of Visual Information in Chinese Characters: Evaluating Large Models' Ability to Recognize and Utilize Radicals
Xiaofeng Wu, Karl Stratos, Wei Xu
The glyphic writing system of Chinese incorporates information-rich visual features in each character, such as radicals that provide hints about meaning or pronunciation. However,…
Model Editing by Standard Fine-Tuning
Govind Gangadhar, Karl Stratos
Standard fine-tuning is considered not as effective as specialized methods for model editing due to its comparatively poor performance. However, it is simple, agnostic to the archi…