4 papers
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
Garvin Guo, Donglei Yu, Yu Chen +6
Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interpretation can be premature: a t…
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
Garvin Guo, Yu Chen, Xiang Wang +4
Recent latent visual reasoning methods achieve substantial gains by inserting continuous latent tokens into multimodal language models. These gains are commonly attributed to the t…
ALIVE: Awakening LLM Reasoning via Adversarial Learning and Instructive Verbal Evaluation
Yiwen Duan, Jing Ye, Xinpei Zhao
The quest for expert-level reasoning in Large Language Models (LLMs) has been hampered by a persistent \textit{reward bottleneck}: traditional reinforcement learning (RL) relies on…
Listening to the Echo: User-Reaction Aware Policy Optimization via Scalar-Verbal Hybrid Reinforcement Learning
Jing Ye, Xinpei Zhao, Lu Xiang +2
While current emotional support dialogue systems typically rely on expert-defined scalar rewards for alignment, these signals suffer from severe information sparsity. They cannot e…