3 papers
cs.CV2026
COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection
Darryl Cherian Jacob, Xinyu Liu, Kai Wang +1
Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM-based VAD methods suff…
cs.CV2025
IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
Jiayi Guo, Chuanhao Yan, Xingqian Xu +4
Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-qu…
cs.CV2025
Slow-Fast Architecture for Video Multi-Modal Large Language Models
Min Shi, Shihao Wang, Chieh-Yun Chen +6
Balancing temporal resolution and spatial detail under limited compute budget remains a key challenge for video-based multi-modal large language models (MLLMs). Existing methods ty…