8 papers
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent
Sudipta Paul, Vijay Srinivasan, Vivek Kulkarni +4
Existing search-augmented LLM agents are trained using Reinforcement Learning to boost its reasoning capabilities. However, these approaches primarily rely on outcome-level rewards…
Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs
Mingyu Wang, Weilin Jin, Wenbo Li +5
Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is inconsistent with or unsupported…
HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models
Weilin Jin, Mingyu Wang, Wenbo Li +5
Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucinations, where model outputs…
Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models
Mingyu Wang, Weilin Jin, Wenbo Li +3
Despite remarkable progress in visual understanding, Multimodal Large Language Models (MLLMs) remain prone to hallucinations when reasoning about spatial relationships, often produ…
PathRouter: Aligning Rewards with Retrieval Quality in Agentic Graph Retrieval-Augmented Generation
Bo Wang, Heyan Huang, Yaolin Li +6
Agentic GraphRAG trains language-model agents to iteratively retrieve and reason over graph-structured evidence, enabling more accurate and context-aware decision-making by efficie…
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Hongcheng Gao, Hailong Qu, Jingyi Tang +18
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predomin…