5 papers
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
Tong Wei, Yijun Yang, Changhao Zhang +4
Multi-turn reinforcement learning (RL) for multi-modal agents built upon vision-language models (VLMs) is hampered by sparse rewards and long-horizon credit assignment. Recent meth…
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
Tong Wei, Yijun Yang, Junliang Xing +3
Reinforcement learning with verifiable outcome rewards (RLVR) has effectively scaled up chain-of-thought (CoT) reasoning in large language models (LLMs). Yet, its efficacy in train…
DeCoDe: Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models
Chengbo He, Bochao Zou, Junliang Xing +3
In human-AI collaboration, a central challenge is deciding whether the AI should handle a task, be deferred to a human expert, or be addressed through collaborative effort. Existin…
Exploring Reliable PPG Authentication on Smartwatches in Daily Scenarios
Jiankai Tang, Jiacheng Liu, Renling Tong +6
Photoplethysmography (PPG) Sensors, widely deployed in smartwatches, offer a simple and non-invasive authentication approach for daily use. However, PPG authentication faces reliab…
Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents
Chengbo He, Bochao Zou, Xin Li +3
Agents have demonstrated their potential in scientific reasoning tasks through large language models. However, they often face challenges such as insufficient accuracy and degenera…