9 papers
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
Yubo Jiang, Xin Yang, Abudukelimu Wuerkaixi +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination--generating content that contradicts visual reality--due to an over-reliance on linguistic priors. W…
Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding
Yubo Jiang, Yitong An, Xin Yang +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination, generating content that contradicts visual reality, due to an over-reliance on linguistic priors. W…
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization
Yubo Jiang, Yitong An, Xin Yang +7
We introduce V-tableR1, a process-supervised reinforcement learning framework that elicits rigorous, verifiable reasoning from multimodal large language models (MLLMs). Current MLL…
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
Jiashu Yang, Chi Zhang, Abudukelimu Wuerkaixi +5
Multimodal document question answering requires retrieving dispersed evidence from visually rich long documents and performing reliable reasoning over heterogeneous information. Ex…
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
Xin Yang, Letian Li, Abudukelimu Wuerkaixi +5
Large language models (LLMs) have demonstrated remarkable and steadily improving performance across a wide range of tasks. However, LLM performance may be highly sensitive to promp…
Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems
Xuxin Cheng, Ke Zeng, Zhiquan Cao +65
Enhancing customer experience is essential for business success, particularly as service demands grow in scale and complexity. Generative artificial intelligence and Large Language…