6 papers
Action-Prior Denoising for Smooth Real-Time Chunking
Dongyang Liu, Zhaowen Zheng, Yu Sun +3
Real-time chunking (RTC) lets chunked action policies operate under inference delay by conditioning a newly generated action chunk on actions already committed by the previous chun…
Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving?
Long Zhang, Yuchen Xia, Bingqing Wei +4
The advent of Large Multimodal Models (LMMs) offers a promising technology to tackle the limitations of modular design in autonomous driving, which often falters in open-world scen…
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
Zijun Min, Bingshuai Liu, Ante Wang +4
Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus…
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
Anxiang Zeng, Haibo Zhang, Hailing Zhang +13
We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…
STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
Yuhan Chen, Yuxuan Liu, Long Zhang +3
Multi-turn interaction remains challenging for online reinforcement learning. A common solution is trajectory-level optimization, which treats each trajectory as a single training…
Compass-Thinker-7B Technical Report
Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6
Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…