2 papers
cs.CV2026
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning
Chen Zhao, Jiajun Ma, Qilong Huang +6
While Multimodal Large Language Models (MLLMs) have advanced video understanding, achieving precise temporal and cross-modal alignment in audiovisual video captioning remains a for…
cs.CV2025
LongCat-Video Technical Report
Meituan LongCat Team, Xunliang Cai, Qilong Huang +8
Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational vid…