Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
MAGE: Modality-Agnostic Music Generation and Target-Source Extraction
Muhammad Usama Saleem, Tejasvi Ravi, Tianyu Xu +4
Recent advances in multimodal audio generation have enabled music synthesis from text, visual cues, and other high-level conditions. However, most systems are designed for a single…
cs.SD2026
SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
Pu Wang, Shinji Watanabe, Hugo Van hamme
Parameter-efficient fine-tuning (PEFT) is a scalable approach for adapting large speech foundation models to new domains. While methods such as LoRA and its state-of-the-art varian…