works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CL2026

When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory

Ruizhe Li, Licheng Zhang, Benfeng Xu +3

Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any qu…

cs.CL2026

When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

Yinfeng Wang, Zhiyuan Yao, Zheren Fu +2

Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, w…

cs.CL2026

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

Zijuan Zhao, Zheren Fu, Hou Xia +3

Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are la…

cs.CV2026

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

Zhixiao Zheng, Zheren Fu, Zhiyuan Yao +3

The paper introduces Groc-PO, a preference‑optimization framework that provides stage‑specific supervision for object grounding, contextual grounding, and grounded reasoning in mul…

cs.CV2026

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

Jingjing Zhang, Lei Zhang, Zheren Fu +1

Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) al…

cs.CV2026

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs

Zhiyuan Yao, Zheren Fu, Zhixiao Zheng +3

Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal s…