collaborators

7 papers

cs.CV2025

Efficient Motion-Aware Video MLLM

Zijia Zhao, Yuqi Huo, Tongtian Yue +5

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges…

cs.CV2025

Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Zijia Zhao, Haoyu Lu, Yuqi Huo +6

Video understanding is a crucial next step for multimodal large language models (MLLMs). Various benchmarks are introduced for better evaluating the MLLMs. Nevertheless, current vi…

cs.CL2025

Baichuan-M1: Pushing the Medical Capability of Large Language Models

Bingning Wang, Haizhou Zhao, Huozhi Zhou +39

The current generation of large language models (LLMs) is typically designed for broad, general-purpose applications, while domain-specific LLMs, especially in vertical fields like…

cs.CV2025

Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Yifan Du, Zikang Liu, Yifan Li +7

Recently, slow-thinking reasoning systems, built upon large language models (LLMs), have garnered widespread attention by scaling the thinking time during inference. There is also…

cs.AI2024

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin +23

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpa…

cs.CV2024

Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining

Han Huang, Yuqi Huo, Zijia Zhao +6

Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-tex…