13 papers
MMAG: A Multi-Control Mixed Audio Generation Benchmark
Zihao Zheng, Xuenan Xu, Jiahao Mei +5
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
Xiaoyu Yang, Yifan Yang, Zengrui Jin +5
Self-supervised learning (SSL) has significantly advanced acoustic representation learning. However, most existing models are optimised for either speech or audio event understandi…
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
Yiming Ren, Xuenan Xu, Ziyang Zhang +3
Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched c…
AOT-POT: Adaptive Operator Transformation for Large-Scale PDE Pre-training
Qitan Lv, Hong Wang, Zhongkai Hao +5
Pre-training neural operators on diverse partial differential equation (PDE) datasets has emerged as a promising direction for building general-purpose surrogate models in scientif…
STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
Chen Zhang, Liwei Liu, Jun Tao +6
Scientific time series are central to scientific AI but are typically sparse, highly heterogeneous, and limited in scale, making unified representation learning particularly challe…
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
Zihao Zheng, Wen Wu, Chao Zhang +2
Current Text-to-Speech (TTS) systems typically use separate models for speech-prompted and text-prompted timbre control. While unifying both control signals into a single model is…