collaborators

5 papers

cs.CV2026

Let ViT Speak: Generative Language-Image Pre-training

Yan Fang, Mengcheng Lan, Zilong Huang +7

In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…

cs.CV2026

iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning

Chang-Bin Zhang, Yujie Zhong, Qiang Zhang +1

While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy duri…

cs.CV2026

TextSculptor: Training and Benchmarking Scene Text Editing

Yiheng Lin, Siyu Jiao, Xiaohan Lan +12

Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…

cs.CV2025

Mr. DETR++: Instructive Multi-Route Training for Detection Transformers with Mixture-of-Experts

Chang-Bin Zhang, Yujie Zhong, Kai Han

Existing methods enhance the training of detection transformers by incorporating an auxiliary one-to-many assignment. In this work, we treat the model as a multi-task framework, si…

cs.CV2025

v-CLR: View-Consistent Learning for Open-World Instance Segmentation

Chang-Bin Zhang, Jinhong Ni, Yujie Zhong +1

In this paper, we address the challenging problem of open-world instance segmentation. Existing works have shown that vanilla visual networks are biased toward learning appearance…