From the 1 of 9 linked papers with an AI index.
9 papers
When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
Yinfeng Wang, Zhiyuan Yao, Zheren Fu +2
Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, w…
MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators
Zijuan Zhao, Zheren Fu, Hou Xia +3
Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are la…
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
Zhixiao Zheng, Zheren Fu, Zhiyuan Yao +3
The paper introduces Groc-PO, a preference‑optimization framework that provides stage‑specific supervision for object grounding, contextual grounding, and grounded reasoning in mul…
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Jingjing Zhang, Lei Zhang, Zheren Fu +1
Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) al…
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
Zhiyuan Yao, Zheren Fu, Zhixiao Zheng +3
Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal s…
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
Jingxuan Han, Wei Liu, Mingyang Zhu +8
Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information…