6 papers
Enhancing Diffusion-Based Quantitatively Controllable Image Generation via Matrix-Form EDM and Adaptive Vicinal Training
Xin Ding, Yun Chen, Sen Zhang +5
Continuous Conditional Diffusion Model (CCDM) is a diffusion-based framework designed to generate high-quality images conditioned on continuous regression labels. Although CCDM has…
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
Agus Gunawan, Samuel Teodoro, Yun Chen +3
Recent advancements in diffusion-based text synthesis have demonstrated significant performance in inserting and editing text within images via inpainting. However, despite the pot…
ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark
Yun Chen, Qi Chen, Zheqi Dai +3
Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend o…
Seedream 4.0: Toward Next-generation Multimodal Image Generation
Team Seedream, :, Yunpeng Chen +48
We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image compositi…
CharaConsist: Fine-Grained Consistent Character Generation
Mengyu Wang, Henghui Ding, Jianing Peng +3
In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications. Although a few works have exp…
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
Junqi Zhao, Jinzheng Zhao, Haohe Liu +5
Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by lear…