8 papers
Beyond Normal References: Discriminative Few-Shot Anomaly Detection
Huan Wang, Jun Shen, Jun Yan +1
This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomalous examples are available a…
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
Yunze Liu, Chi-Hao Wu, Enmin Zhou +1
Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) encoders extend this setting…
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
Peiran Wu, Yunze Liu, Chi-Hao Wu +2
Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio v…
FGSVQA: Frequency-Guided Short-form Video Quality Assessment
Xinyi Wang, Angeliki Katsenou, Junxiao Shen +1
Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content variation, and mixed distortio…
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
Peiran Wu, Zhuorui Yu, Yunze Liu +3
The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when e…
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
Xinyi Wang, Angeliki Katsenou, Junxiao Shen +1
The prevalence of user-generated content (UGC) on platforms such as YouTube and TikTok has rendered no-reference (NR) perceptual video quality assessment (VQA) vital for optimizing…