8 citations · 8 across the 5 of their papers we have counts for
11 papers
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
Changyao Tian, Hao Li, Gen Luo +11
Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs th…
Sequential Diffusion Language Models
Yangzhou Liu, Yue Cao, Hao Li +13
Diffusion language models (DLMs) have strong theoretical efficiency but are limited by fixed-length decoding and incompatibility with key-value (KV) caches. Block diffusion mitigat…
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Gen Luo, Wenhan Dou, Wenhao Li +9
This paper focuses on monolithic Multimodal Large Language Models (MLLMs), which integrate visual encoding and language decoding into a single model. Existing structures and pre-tr…
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
Teng Li, Quanfeng Lu, Lirui Zhao +5
Unified image understanding and generation has emerged as a promising paradigm in multimodal artificial intelligence. Despite recent progress, the optimal architectural design for…
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
Chenyu Yang, Shiqian Su, Shi Liu +11
The rapid advancement of large Vision-Language Models (VLMs) has propelled the development of pure-vision-based GUI Agents, capable of perceiving and operating Graphical User Inter…
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
Yan Li, Changyao Tian, Renqiu Xia +7
We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise…