collaborators

7 papers

cs.CV2025

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Sucheng Ren, Qihang Yu, Ju He +2

Diffusion-based Transformers have demonstrated impressive generative capabilities, but their high computational costs hinder practical deployment, for example, generating an $8192\…

cs.CV2025

ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling

Qihao Liu, Ju He, Qihang Yu +2

In recent years, video generation has seen significant advancements. However, challenges still persist in generating complex motions and interactions. To address these challenges,…

cs.CV2025

Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Sucheng Ren, Qihang Yu, Ju He +3

Autoregressive (AR) modeling, known for its next-token prediction paradigm, underpins state-of-the-art language and visual generative models. Traditionally, a ``token'' is treated…

cs.CV2025

FlowTok: Flowing Seamlessly Across Text and Image Tokens

Ju He, Qihang Yu, Qihao Liu +1

Bridging different modalities lies at the heart of cross-modality generation. While conventional approaches treat the text modality as a conditioning signal that gradually guides t…

cs.CV2025

Dictionary-based Framework for Interpretable and Consistent Object Parsing

Tiezheng Zhang, Qihang Yu, Alan Yuille +1

In this work, we present CoCal, an interpretable and consistent object parsing framework based on dictionary-based mask transformer. Designed around Contrastive Components and Logi…

cs.CV2025

Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

Dongwon Kim, Ju He, Qihang Yu +4

Image tokenizers form the foundation of modern text-to-image generative models but are notoriously difficult to train. Furthermore, most existing text-to-image models rely on large…