55 papers
AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval
Haoyu Zuo, Yibo Yan, Xin Zou +4
Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings pe…
Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?
Yubo Gao, Haotian Wu, Xiaoyu Xu +9
Existing methods for multimodal sentiment analysis (MSA) under missing modalities usually follow a repair-first paradigm. We revisit this assumption and ask: \emph{should every mis…
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
Jiahao Huo, Wenjie Qu, Yibo Yan +5
The paper introduces SAMark, a text watermarking method that remains detectable even after paragraph‑level paraphrasing by removing reliance on sentence order and using a hyperboli…
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation
Xin Zou, Haolin Deng, Yibo Yan +5
Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on l…
Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning
Xin Zou, Haolin Deng, Yibo Yan +6
Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cross-view consistency, \texti…
MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework
Haowen Xiang, Yibo Yan, Jiahao Huo +4
Multi-vector visual document retrievers achieve strong fine-grained matching by representing each page with multiple vectors from deep Vision-Language Models (VLMs), but this desig…