6 papers
The Illusion of : Evaluating the Breakdown of Counterfactual Reasoning in LLMs
Yucheng Wang, Yuetian Du, Zhengyi Liu +8
Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks large…
MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG
Qiuyi Qi, Tian Liang, Jiamu Wang +7
Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching and when to answer. Existing RL-based methods rely on external supervisio…
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
Yuetian Du, Yucheng Wang, Zhenyuan Chen +9
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these mo…
Living-Harness Is an Interactive-Agent Evolver
Yuetian Du, Yucheng Wang, He Xu +9
Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedba…
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
Qiuyi Qi, Tian Liang, Mutian Bao +8
Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to traject…
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs
Qiuyi Qi, Jinjian Zhang, Mutian Bao +9
Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their r…