collaborators

6 papers

cs.RO2026

Action-Prior Denoising for Smooth Real-Time Chunking

Dongyang Liu, Zhaowen Zheng, Yu Sun +3

Real-time chunking (RTC) lets chunked action policies operate under inference delay by conditioning a newly generated action chunk on actions already committed by the previous chun…

cs.RO2026

Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving?

Long Zhang, Yuchen Xia, Bingqing Wei +4

The advent of Large Multimodal Models (LMMs) offers a promising technology to tackle the limitations of modular design in autonomous driving, which often falters in open-world scen…

cs.LG2026

Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR

Zijun Min, Bingshuai Liu, Ante Wang +4

Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus…

cs.AI2025

Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE

Anxiang Zeng, Haibo Zhang, Hailing Zhang +13

We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…

cs.AI2025

STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization

Yuhan Chen, Yuxuan Liu, Long Zhang +3

Multi-turn interaction remains challenging for online reinforcement learning. A common solution is trajectory-level optimization, which treats each trajectory as a single training…

cs.AI2025

Compass-Thinker-7B Technical Report

Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6

Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…