collaborators

6 papers

cs.SD2026

DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation

Zhongjie Duan, Shengchuan Gao, Hong Zhang +1

Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Musi…

cs.CV2026

TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

Yuze Sun, Zhongjie Duan, Yingda Chen

Although general text-to-image models excel in open-domain generation, their performance degrades significantly in specialized downstream domains, particularly when generating imag…

cs.CV2026

Compressing Image Style Training into a Single Model Forward

Zhongjie Duan, Yingda Chen

Diffusion-based style transfer must balance inference efficiency with stylization fidelity. Adapter-based methods are efficient, but they inject style as an external condition and…

cs.LG2026

Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion

Zhongjie Duan, Hong Zhang, Yingda Chen

Controllable diffusion methods have substantially expanded the practical utility of diffusion models, but they are typically developed as isolated, backbone-specific systems with i…

cs.CV2025

Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

Hong Zhang, Zhongjie Duan, Xingjun Wang +6

Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly i…

cs.CV2025

EliGen: Entity-Level Controlled Image Generation with Regional Attention

Hong Zhang, Zhongjie Duan, Xingjun Wang +2

Recent advancements in diffusion models have significantly advanced text-to-image generation, yet global text prompts alone remain insufficient for achieving fine-grained control o…