3 papers
cs.CV2026
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
Chen Ling, Hanqian Li, Dongnan Liu +9
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models t…
cs.AI2026
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
Qiuyi Qi, Tian Liang, Mutian Bao +8
Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to traject…
cs.AI2026
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs
Qiuyi Qi, Jinjian Zhang, Mutian Bao +9
Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their r…