4 papers
WAT: Online Video Understanding Needs Watching Before Thinking
Zifan Han, Hongbo Sun, Jinglin Xu +6
Multimodal Large Language Models (MLLMs) have shown strong capabilities in image understanding, motivating recent efforts to extend them to video reasoning. However, existing Video…
HunyuanOCR Technical Report
Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan +29
This paper presents HunyuanOCR, a commercial-grade, open-source, and lightweight (1B parameters) Vision-Language Model (VLM) dedicated to OCR tasks. The architecture comprises a Na…
Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection
Canhui Tang, Sanping Zhou, Haoyue Shi +1
Zero-Shot Video Anomaly Detection (ZS-VAD) requires temporally localizing anomalies without target domain training data, which is a crucial task due to various practical concerns,…
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
Canhui Tang, Zifan Han, Hongbo Sun +7
Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs.…