activity
20242026
collaborators

9 papers

cs.LG2026

Hide to Guide: Learning via Semantic Masking

Ruitao Liu, Qinghao Hu, Alex Hu +6

Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limi…

cs.LG2026

Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter

Qinghao Hu, Shang Yang, Junxian Guo +7

The emergence of Large Language Models (LLMs) with strong reasoning capabilities marks a significant milestone, unlocking new frontiers in complex problem-solving. However, trainin…

cs.LG2026

Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs

Luke J. Huang, Zhuoyang Zhang, Qinghao Hu +2

Asynchronous reinforcement learning has become increasingly central to scaling LLM post-training, delivering major throughput gains by decoupling rollout generation from policy upd…

cs.RO2026

ForeAct: Steering Your VLA with Efficient Visual Foresight Planning

Zhuoyang Zhang, Shang Yang, Qinghao Hu +5

Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We…

cs.CV2025

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi +11

We introduce a full-stack framework that scales up reasoning in vision-language models (VLMs) to long videos, leveraging reinforcement learning. We address the unique challenges of…

cs.CL2025

Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search

Yuxian Gu, Qinghao Hu, Shang Yang +4

We present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving g…