activity
20242026
collaborators

6 papers

cs.LG2026

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

Qing Miao, Yiming Zhao, Jing Yang +5

Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models (LLMs), yet it remains limit…

cs.IR2026

Tiny-Critic RAG: Empowering Agentic Fallback with Parameter-Efficient Small Language Models

Yichao Wu, Penghao Liang, Yafei Xiang +5

Retrieval-Augmented Generation (RAG) grounds Large Language Models (LLMs) to mitigate factual hallucinations. Recent paradigms shift from static pipelines to Modular and Agentic RA…

cs.CV2025

Pretrained Reversible Generation as Unsupervised Visual Representation Learning

Rongkun Xue, Jinouwen Zhang, Yazhe Niu +4

Recent generative models based on score matching and flow matching have significantly advanced generation tasks, but their potential in discriminative tasks remains underexplored.…

cs.SD2025

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

Rongkun Xue, Yazhe Niu, Shuai Hu +3

Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantiz…

cs.AI2024

ReZero: Boosting MCTS-based Algorithms by Backward-view and Entire-buffer Reanalyze

Chunyu Xuan, Yazhe Niu, Yuan Pu +3

Monte Carlo Tree Search (MCTS)-based algorithms, such as MuZero and its derivatives, have achieved widespread success in various decision-making domains. These algorithms employ th…

cs.LG2024

Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective

Jinouwen Zhang, Rongkun Xue, Yazhe Niu +4

Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement le…