activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

Efficient Motion-Aware Video MLLM

Zijia Zhao, Yuqi Huo, Tongtian Yue +5

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges…

cs.CV2025

Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Zijia Zhao, Haoyu Lu, Yuqi Huo +6

Video understanding is a crucial next step for multimodal large language models (MLLMs). Various benchmarks are introduced for better evaluating the MLLMs. Nevertheless, current vi…

cs.CV2025

Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Yifan Du, Zikang Liu, Yifan Li +7

Recently, slow-thinking reasoning systems, built upon large language models (LLMs), have garnered widespread attention by scaling the thinking time during inference. There is also…

cs.CV2024

Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining

Han Huang, Yuqi Huo, Zijia Zhao +6

Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-tex…

cs.CV2024

Exploring the Design Space of Visual Context Representation in Video MLLMs

Yifan Du, Yuqi Huo, Kun Zhou +7

Video Multimodal Large Language Models (MLLMs) have shown remarkable capability of understanding the video semantics on various downstream tasks. Despite the advancements, there is…

cs.CV2024

Towards Event-oriented Long Video Understanding

Yifan Du, Kun Zhou, Yuqi Huo +7

With the rapid development of video Multimodal Large Language Models (MLLMs), numerous benchmarks have been proposed to assess their video understanding capability. However, due to…