1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.AI2026
daVinci-LLM:Towards the Science of Pretraining
Yiwei Qin, Yixiu Liu, Tiantian Mi +12
The foundational pretraining phase determines a model's capability ceiling, as post-training struggles to overcome capability foundations established during pretraining, yet it rem…
cs.AI2026
Data Darwinism Part II: DataEvolve -- AI can Autonomously Evolve Pretraining Data Curation
Tiantian Mi, Dongming Shan, Zhen Huang +6
Data Darwinism (Part I) established a ten-level hierarchy for data processing, showing that stronger processing can unlock greater data value. However, that work relied on manually…
cs.LG2024★ 1 cited
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
Yuhong Chou, Man Yao, Kexin Wang +7
Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax a…