2 papers
cs.CV2026
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
Zheng Wang, Haoran Chen, Haoxuan Qin +3
Long video understanding is challenging due to dense visual redundancy, long-range temporal dependencies, and the tendency of chain-of-thought and retrieval-based agents to accumul…
cs.CV2025
A Benchmark for Crime Surveillance Video Analysis with Large Models
Haoran Chen, Dong Yi, Moyan Cao +3
Anomaly analysis in surveillance videos is a crucial topic in computer vision. In recent years, multimodal large language models (MLLMs) have outperformed task-specific models in v…