activity
20242026
collaborators

7 papers

cs.CV2026

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework

Guanzhou Chen, Erfei Cui, Changyao Tian +6

Instruction-based image editing has emerged as a key capability for unified multimodal models (UMMs), yet constructing large-scale, diverse, and high-quality editing datasets witho…

cs.CV2026

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

Changyao Tian, Danni Yang, Guanzhou Chen +26

Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade-offs between maintaining strong semantic comprehension and acq…

cs.CV2025

Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space

Yan Li, Changyao Tian, Renqiu Xia +7

We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise…

cs.CV2025

NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints

Changyao Tian, Hao Li, Gen Luo +11

Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs th…

cs.CV2025

Docopilot: Improving Multimodal Models for Document-Level Understanding

Yuchen Duan, Zhe Chen, Yusong Hu +9

Despite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lac…

cs.CV2025

Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

Yuchen Duan, Weiyun Wang, Zhe Chen +7

Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processi…