collaborators

5 papers

cs.CV2026

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Wujian Peng, Lingchen Meng, Yuxuan Cai +7

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokeni…

cs.RO2026

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Qiuyue Wang, Mingsheng Li, Jian Guan +37

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generali…

cs.CV2026

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

Zhuohan Liu, Wujian Peng, Yitong Chen +1

Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bindings, object relationships…

cs.CV2026

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

Wujian Peng, Lingchen Meng, Yitong Chen +7

Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a…

cs.CV2025

Enhancing Vision Foundation Models via Multimodal Continual Pre-Training

Yitong Chen, Lingchen Meng, Wujian Peng +4

Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowi…