most citedS2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

1 citations · 1 across the 6 of their papers we have counts for

collaborators

10 papers

eess.AS2026

Preference Optimization for Non-Verbal Vocalization Synthesis

Haoyang Li, Chenglin Xu, Junchuan Zhao +4

Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains po…

eess.AS2026

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

Dehui Gao, Zhixian Zhao, Zhennan Lin +14

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentri…

cs.SD2026

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

Zeyue Tian, Lei Ke, Zhaoyang Liu +8

Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework,…

cs.SD2026

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech

Wei Xue, Junlan Feng, Shilei Zhang +9

Recent advances in text-to-speech (TTS) have greatly improved speech naturalness, speaker similarity, and controllability. However, most existing controllable TTS systems still rel…

cs.SD2026

Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound

Liumeng Xue, Ziya Zhou, Jiahao Pan +20

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and gene…

cs.CL20261 cited

S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

Feng Jiang, Zhiyu Lin, Yiyang Liu +6

Recent advances in large language models (LLMs) have fundamentally reshaped speech-to-speech (S2S) systems, enabling increasingly natural spoken interaction. However, existing benc…