From the 1 of 11 linked papers with an AI index.
11 papers
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
Ying Zhang, Yu Zhao, Xuhui Sui +5
The paper introduces a federated learning framework for multimodal knowledge graph completion that recovers missing multimodal information with a hyper-modal imputation diffusion e…
R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
Yibo Wang, Haotian Luo, Huanjin Yao +8
Chain-of-Thought (CoT) reasoning enhances large language models (LLMs) by enabling step-by-step problem-solving, yet its extension to Long-CoT introduces substantial computational…
Effective Policy Learning for Multi-Agent Online Coordination Beyond Submodular Objectives
Qixin Zhang, Yan Sun, Can Jin +5
In this paper, we present two effective policy learning algorithms for multi-agent online coordination(MA-OC) problem. The first one, \texttt{MA-SPL}, not only can achieve the opti…
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
Botai Yuan, Yutian Zhou, Yingjie Wang +9
Recent benchmarks for medical Large Vision-Language Models (LVLMs) emphasize leaderboard accuracy, overlooking reliability and safety. We study sycophancy -- models' tendency to un…
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
Rong-Cheng Tu, Wenhao Sun, Hanzhe You +4
Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a compositional query, consisting of a reference image and a modifying text-without relying on anno…
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
Rong-Cheng Tu, Zhao Jin, Jingyi Liao +4
Existing Zero-Shot Composed Image Retrieval (ZS-CIR) methods typically train adapters that convert reference images into pseudo-text tokens, which are concatenated with the modifyi…