7 papers · 1 filter
ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory
Shidu Ren, Yunze Liu, Xing Liu +3
Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associate…
Beyond Normal References: Discriminative Few-Shot Anomaly Detection
Huan Wang, Jun Shen, Jun Yan +1
This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomalous examples are available a…
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
Yunze Liu, Chi-Hao Wu, Enmin Zhou +1
Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) encoders extend this setting…
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
Peiran Wu, Yunze Liu, Chi-Hao Wu +2
Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio v…
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
Peiran Wu, Zhuorui Yu, Yunze Liu +3
The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when e…
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
Peiran Wu, Yunze Liu, Miao Liu +1
Humans excel at spatial-temporal reasoning, effortlessly interpreting dynamic visual events from an egocentric viewpoint. However, whether multimodal large language models (MLLMs)…