3 papers
cs.SD2026
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training
Shun Lei, Huaicheng Zhang, Dapeng Wu +8
Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based…
eess.AS2026
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
Dapeng Wu, Shun Lei, Wei Tan +5
Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture mu…
cs.CV2026
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
Yun Wang, Long Zhang, Jingren Liu +9
Video Large Language Models (Video-LLMs) have shown strong video understanding, yet their application to long-form videos remains constrained by limited context windows. A common w…