6 papers
ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning
Qing Miao, Yiming Zhao, Jing Yang +5
Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models (LLMs), yet it remains limit…
Tiny-Critic RAG: Empowering Agentic Fallback with Parameter-Efficient Small Language Models
Yichao Wu, Penghao Liang, Yafei Xiang +5
Retrieval-Augmented Generation (RAG) grounds Large Language Models (LLMs) to mitigate factual hallucinations. Recent paradigms shift from static pipelines to Modular and Agentic RA…
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
Rongkun Xue, Jinouwen Zhang, Yazhe Niu +4
Recent generative models based on score matching and flow matching have significantly advanced generation tasks, but their potential in discriminative tasks remains underexplored.…
HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
Rongkun Xue, Yazhe Niu, Shuai Hu +3
Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantiz…
ReZero: Boosting MCTS-based Algorithms by Backward-view and Entire-buffer Reanalyze
Chunyu Xuan, Yazhe Niu, Yuan Pu +3
Monte Carlo Tree Search (MCTS)-based algorithms, such as MuZero and its derivatives, have achieved widespread success in various decision-making domains. These algorithms employ th…
Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
Jinouwen Zhang, Rongkun Xue, Yazhe Niu +4
Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement le…