From the 1 of 13 linked papers with an AI index.
13 papers
When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory
Ruizhe Li, Licheng Zhang, Benfeng Xu +3
Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any qu…
When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
Yinfeng Wang, Zhiyuan Yao, Zheren Fu +2
Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, w…
MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators
Zijuan Zhao, Zheren Fu, Hou Xia +3
Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are la…
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
Zhixiao Zheng, Zheren Fu, Zhiyuan Yao +3
The paper introduces Groc-PO, a preference‑optimization framework that provides stage‑specific supervision for object grounding, contextual grounding, and grounded reasoning in mul…
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Jingjing Zhang, Lei Zhang, Zheren Fu +1
Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) al…
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
Zhiyuan Yao, Zheren Fu, Zhixiao Zheng +3
Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal s…