6 papers
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current…
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal gro…
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
Chenting Wang, Yuhan Zhu, Yicheng Xu +6
Large-scale video-text pretraining achieves strong performance but depends on noisy, synthetic captions with limited semantic coverage, often overlooking implicit world knowledge s…
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
Chunxu Liu, Jiyuan Yang, Ruopeng Gao +4
Multimodal embeddings are widely used in downstream tasks such as multimodal retrieval, enabling alignment of interleaved modalities in a shared representation space. While recent…
Motion-Aware Generative Frame Interpolation
Guozhen Zhang, Yuhan Zhu, Yutao Cui +3
Flow-based frame interpolation methods ensure motion stability through estimated intermediate flow but often introduce severe artifacts in complex motion regions. Recent generative…
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Xinhao Li, Yi Wang, Jiashuo Yu +10
Long-context video modeling is critical for multimodal large language models (MLLMs), enabling them to process movies, online video streams, and so on. Despite its advances, handli…