works on

From the 1 of 41 linked papers with an AI index.

activity
20242026
most citedSLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

3 citations · 5 across the 22 of their papers we have counts for

collaborators
Showing 2026 · cs.SDShow all

8 papers · 2 filters

cs.SD2026

VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition

Yukun Chen, Tianrui Wang, Zhaoxi Mu +2

The paper introduces VocalRender, a system that can directly synthesize singing voices from musical scores—including lyrics, pitches, note values, and tempo—without needing separat…

cs.SD2026

QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection

Duc-Tuan Truong, Tianchi Liu, Ruijie Tao +3

Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-ce…

cs.SD2026

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

Yue Heng Yeo, Haoyang Li, Yizhou Peng +6

Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augme…

cs.SD2026

MMAE: A Massive Multitask Audio Editing Benchmark

Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…

cs.SD2026

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

Kai Li, Can Shen, Yile Liu +31

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks ar…

cs.SD2026

Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input

Changsong Liu, Tianrui Wang, Ye Ni +2

Streaming TTS that receives streaming text is essential for interactive systems, yet this scheme faces two major challenges: unnatural prosody due to missing lookahead and long-for…