collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

Shidu Ren, Yunze Liu, Xing Liu +3

Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associate…

cs.CV2026

Beyond Normal References: Discriminative Few-Shot Anomaly Detection

Huan Wang, Jun Shen, Jun Yan +1

This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomalous examples are available a…

cs.CV2026

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

Yunze Liu, Chi-Hao Wu, Enmin Zhou +1

Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) encoders extend this setting…

cs.CV2026

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

Peiran Wu, Yunze Liu, Chi-Hao Wu +2

Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio v…

cs.CV2025

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

Peiran Wu, Zhuorui Yu, Yunze Liu +3

The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when e…

cs.CV2025

ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Peiran Wu, Yunze Liu, Miao Liu +1

Humans excel at spatial-temporal reasoning, effortlessly interpreting dynamic visual events from an egocentric viewpoint. However, whether multimodal large language models (MLLMs)…