From the 2 of 8 linked papers with an AI index.
8 papers
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
Yu Zhao, Ying Zhang, Xuhui Sui +4
The paper introduces a teacher‑student framework called Hindsight Distillation (HinD) that uses privileged answer information to generate reasoning trajectories for a multimodal LL…
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
Ying Zhang, Yu Zhao, Xuhui Sui +5
The paper introduces a federated learning framework for multimodal knowledge graph completion that recovers missing multimodal information with a hyper-modal imputation diffusion e…
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao +9
Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate r…
ClueAegis: Heuristic-to-Reasoning Cognitive-skill Learning for Unified Evidence-based Synthetic Image Detection
Huangsen Cao, Hongkang Chu, Yuxi Li +6
The rapid advancement of generative models has made synthetic images increasingly realistic, challenging reliable detection. Existing methods are often limited to end-to-end classi…
Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction
Baohang Zhou, Kehui Song, Rize Jin +5
Multimodal information extraction (MIE) constitutes a set of essential tasks aimed at extracting structural information from Web texts with integrating images, to facilitate the st…
Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
Xinying Qian, Ying Zhang, Yu Zhao +3
Temporal Knowledge Graph Question Answering (TKGQA) aims to answer time-sensitive questions by leveraging factual information from Temporal Knowledge Graphs (TKGs). While previous…