25 papers
AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval
Haoyu Zuo, Yibo Yan, Xin Zou +4
Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings pe…
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation
Xin Zou, Haolin Deng, Yibo Yan +5
Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on l…
Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning
Xin Zou, Haolin Deng, Yibo Yan +6
Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cross-view consistency, \texti…
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
Haolin Deng, Xin Zou, Zhiwei Jin +3
Multimodal hallucination remains a persistent challenge for Vision-Language Models (VLMs). Standard textual Direct Preference Optimization (DPO) often fails to mitigate it due to a…
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
Shu-Hao Zhang, Le-Tong Huang, Xiang-Sheng Deng +5
Quantization has emerged as a mainstream approach for deploying Large Language Models (LLMs) on resource-constrained devices, yet compressing precision below 4-bit typically causes…
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
Fanpu Cao, Xin Zou, Xuming Hu +1
Multimodal large language models (MLLMs) have become a key interface for visual reasoning and grounded question answering, yet they remain vulnerable to visual hallucinations, wher…