1 paper · 1 filter
Changdi Yang, Fengquan Jiao, Haochih Lin +14
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight…