Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
Zhepei Wei, Xinyu Zhu, Wei-Lin Chen +3
Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the underlying geometry of the res…
cs.LG2026
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
Yusong Wu, Stephen Brade, Aleksandra Teng Ma +6
Most applications of generative AI involve a sequential interaction in which a person inputs a prompt and waits for a response, and where reaction time and adaptivity are not impor…