2 papers
cs.LG2026
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
Zhangyang Yao, Haiyan Zhao, Haoyu Wang +3
Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive modules. However, automating this allocat…
cs.CL2025
Yi-Lightning Technical Report
Alan Wake, Bei Chen, C. X. Lv +41
This technical report presents Yi-Lightning, our latest flagship large language model (LLM). It achieves exceptional performance, ranking 6th overall on Chatbot Arena, with particu…