most citedInternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

8 citations · 8 across the 5 of their papers we have counts for

collaborators

11 papers

cs.CV2025

NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints

Changyao Tian, Hao Li, Gen Luo +11

Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs th…

cs.CL2025

Sequential Diffusion Language Models

Yangzhou Liu, Yue Cao, Hao Li +13

Diffusion language models (DLMs) have strong theoretical efficiency but are limited by fixed-length decoding and incompatibility with key-value (KV) caches. Block diffusion mitigat…

cs.CV2025

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Gen Luo, Wenhan Dou, Wenhao Li +9

This paper focuses on monolithic Multimodal Large Language Models (MLLMs), which integrate visual encoding and language decoding into a single model. Existing structures and pre-tr…

cs.CV2025

UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Teng Li, Quanfeng Lu, Lirui Zhao +5

Unified image understanding and generation has emerged as a promising paradigm in multimodal artificial intelligence. Despite recent progress, the optimal architectural design for…

cs.AI2025

ZeroGUI: Automating Online GUI Learning at Zero Human Cost

Chenyu Yang, Shiqian Su, Shi Liu +11

The rapid advancement of large Vision-Language Models (VLMs) has propelled the development of pure-vision-based GUI Agents, capable of perceiving and operating Graphical User Inter…

cs.CV2025

Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space

Yan Li, Changyao Tian, Renqiu Xia +7

We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise…