works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24

VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…

cs.CV2026

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

Ming Xie, Zizheng Huang, Xudong Tan +6

While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…

cs.CV2026

Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Xiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu +12

Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse y…

cs.CV2025

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

Jiashuo Yu, Yue Wu, Meng Chu +14

We present VRBench, the first long narrative video benchmark crafted for evaluating large models' multi-step reasoning capabilities, addressing limitations in existing evaluations…

cs.CV2025

Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training

Zizheng Huang, Haoxing Chen, Jiaqi Li +4

Recent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies…

cs.CV2025

Efficient Transfer Learning for Video-language Foundation Models

Haoxing Chen, Zizheng Huang, Yan Hong +5

Pre-trained vision-language models provide a robust foundation for efficient transfer learning across various downstream tasks. In the field of video action recognition, mainstream…