5 papers
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
Yunze Liu, Chi-Hao Wu, Enmin Zhou +1
Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) encoders extend this setting…
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
Peiran Wu, Yunze Liu, Chi-Hao Wu +2
Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio v…
Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
Jiazheng Li, Chi-Hao Wu, Yunze Liu +3
Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even…
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
Xinyi Zheng, Yunze Liu, Chi-Hao Wu +5
We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaff…
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
Peiran Wu, Zhuorui Yu, Yunze Liu +3
The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when e…