collaborators

5 papers

cs.CV2026

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Shangwen Zhu, Qianyu Peng, Zhao Pu +12

Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace th…

cs.CV2026

Accelerating Diffusion Sampling via Exploiting Local Transition Coherence

Shangwen Zhu, Han Zhang, Zhantao Yang +4

Text-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the lengthy sampling time of the de…

cs.CV2026

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

Siyong Jian, Siyuan Li, Luyuan Zhang +5

Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the policy while keeping the VQ dec…

cs.CV2025

Addressing the ID-Matching Challenge in Long Video Captioning

Zhantao Yang, Huangji Wang, Ruili Feng +6

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal u…

cs.CV2025

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan +13

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…