most citedContextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 22 of their papers we have counts for

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Zixuan Huang, Yang Zhou, Kaixuan Wang +7

Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision…

cs.AI2026

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

Zhongyi Li, Wan Tian, Jinju Chen +4

Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such systems is limited by credit assi…

cs.AI2026

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

Zehao Chen, Gongxun Li, Tianxiang Ai +9

As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow highly confident, further training…

cs.AI2026

Does Your Reasoning Model Implicitly Know When to Stop Thinking?

Zixuan Huang, Xin Xia, Yuxi Ren +11

Recent advancements in large reasoning models (LRMs) have greatly improved their capabilities on complex reasoning tasks through Long Chains of Thought (CoTs). However, this approa…

cs.AI2026

Real-Time Aligned Reward Model beyond Semantics

Zixuan Huang, Xin Xia, Yuxi Ren +10

Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for aligning large language models (LLMs) with human preferences, yet it is susceptible to reward overoptim…

cs.AI2026

Adaptive Robust Estimator for Multi-Agent Reinforcement Learning

Zhongyi Li, Wan Tian, Jingyu Chen +8

Multi-agent collaboration has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models, yet it suffers from interaction-level ambiguity that…