From the 2 of 16 linked papers with an AI index.
16 papers
MMFGU: Multimodal Federated Graph Unlearning
Haodong Lu, Zekai Chen, Weiwei Ji +5
Multimodal federated graph learning enables clients to collaboratively train graph models over structural, textual, and visual signals without sharing private local data. However,…
FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning
Zekai Chen, Haodong Lu, Shihao Li +5
The paper introduces FedOGL, a framework for federated multimodal graph learning that mitigates catastrophic forgetting by preserving semantic and structural memory through client-…
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Xinyu Liu, Shihao Li, Weihong Lin +10
The paper introduces ReBind, a framework that uses structured instructions with explicit reference tokens to improve multi‑reference image‑conditioned video editing, enabling preci…
CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales
Xinlong Chen, Jiafu Tang, Yue Ding +12
Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can ob…
CoVEBench: Can Video Editing Models Handle Complex Instructions?
Jiangtao Wu, Jiaming Wang, Yiwen He +7
While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional. A single prompt…
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
Jiahao Wang, An Ping, Yanghai Wang +13
While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…