4 papers
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
Zongcai Du, Guilin Deng, Xiaofeng Guo +8
Recent progress in diffusion-based Singing Voice Synthesis (SVS) demonstrates strong expressiveness but remains limited by data scarcity and model scalability. We introduce a two-s…
Yan: Foundational Interactive Video Generation
Deheng Ye, Fangyun Zhou, Jiacheng Lv +15
We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three…
PRJ: Perception-Retrieval-Judgement for Generated Images
Qiang Fu, Zonglei Jing, Zonghao Ying +1
The rapid progress of generative AI has enabled remarkable creative capabilities, yet it also raises urgent concerns regarding the safety of AI-generated visual content in real-wor…
Playable Game Generation
Mingyu Yang, Junyou Li, Zhongbin Fang +5
In recent years, Artificial Intelligence Generated Content (AIGC) has advanced from text-to-image generation to text-to-video and multimodal video synthesis. However, generating pl…