activity
20242026
collaborators

6 papers

cs.CV2026

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

Jie Zhu, Mingyu Ding, Boqiang Duan +2

Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improve its performance, the under…

cs.CV2026

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

Tengfei Liu, Yang Shi, Xuanyu Zhu +17

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing b…

cs.CV2026

A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction

Jie Zhu, Hanghang Ma, Jia Wang +6

In this work, we introduce Wallaroo, a simple autoregressive baseline that leverages next-token prediction to unify multi-modal understanding, image generation, and editing at the…

cs.CV2025

Auditing Data Provenance in Real-world Text-to-Image Diffusion Models for Privacy and Copyright Protection

Jie Zhu, Leye Wang

Text-to-image diffusion model since its propose has significantly influenced the content creation due to its impressive generation capability. However, this capability depends on l…

cs.CV2025

A Unified and Scalable Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability

Jie Zhu, Jirong Zha, Ding Li +1

Self-supervised learning shows promise in harnessing extensive unlabeled data, but it also confronts significant privacy concerns, especially in vision. In this paper, we perform m…

cs.CV2024

MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts

Jie Zhu, Yixiong Chen, Mingyu Ding +3

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particul…