collaborators

13 papers

cs.CV2026

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

Yunze Liu, Chi-Hao Wu, Enmin Zhou +1

Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) encoders extend this setting…

cs.CV2026

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

Peiran Wu, Yunze Liu, Chi-Hao Wu +2

Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio v…

cs.CV2026

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

Jiazheng Li, Chi-Hao Wu, Yunze Liu +3

Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even…

cs.CV2026

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)

Guanyi Qin, Jie Liang, Bingbing Zhang +50

In this paper, we present an overview of the NTIRE 2026 challenge on the 3rd Restore Any Image Model in the Wild, specifically focusing on Track 1: Professional Image Quality Asses…

cs.CV2026

SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA

Xinyi Zheng, Yunze Liu, Chi-Hao Wu +5

We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaff…

cs.CV2026

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

Peiran Wu, Zhuorui Yu, Yunze Liu +3

The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when e…