4 papers
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
Xu Guo, Fulong Ye, Qichao Sun +7
Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-base…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
Qi He, Gus Xia, Ziyu Wang
Hierarchical planning is a powerful approach to model long sequences structurally. Aside from considering hierarchies in the temporal structure of music, this paper explores an eve…
Advancing Multi-Party Dialogue Framework with Speaker-ware Contrastive Learning
Zhongtian Hu, Qi He, Ronghan Li +2
Multi-party dialogues, common in collaborative scenarios like brainstorming sessions and negotiations, pose significant challenges due to their complexity and diverse speaker roles…