Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma
Chongyu Fan, Pengfei Liu, Jingjia Huang +2
Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs). To achieve sample efficiency, modern…
cs.LG2025
Virtual Width Networks
Seed, Baisheng Li, Banggu Wu +115
We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN d…