2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2026★ 2 cited
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
cs.SD2026
Audio ControlNet for Fine-Grained Audio Generation and Editing
Haina Zhu, Yao Xiao, Xiquan Li +5
We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over at…
cs.CV2026
HY3D-Bench: Generation of 3D Assets
Team Hunyuan3D, :, Bowen Zhang +22
While recent advances in neural representations and generative models have revolutionized 3D content creation, the field remains constrained by significant data processing bottlene…