2 papers
cs.LG2026
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
Zhangyi Hu, Chenhui Liu, Tian Huang +6
Recently, Reinforcement Learning with Verifiable Rewards (RLVR) and Test-Time Scaling (TTS) have advanced LLM code generation through executable verification. Yet Ground-Truth Unit…
cs.LG2026
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
Zhangyang Yao, Haiyan Zhao, Haoyu Wang +3
Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive modules. However, automating this allocat…