6 papers
A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
Siyi Wang, James Bailey, Ting Dang
While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly un…
CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
Siyi Wang, Shihong Tan, Siyi Liu +4
Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In cont…
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
Hongyu Jin, Siyi Wang, Yang Xiao +10
Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory. Despite recent progress in large…
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
Yang Xiao, Siyi Wang, Eun-Jung Holden +1
Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fra…
Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory
Yang Xiao, Siyi Wang, Han Yin +4
Large audio language models (LALMs) process both speech and environmental acoustic cues, yet struggle to retain non-speech information across multi-turn interactions. The performan…
Emotion-Aware Quantization for Discrete Speech Representations: An Analysis of Emotion Preservation
Haoguang Zhou, Siyi Wang, Jingyao Wu +2
Modern speech systems increasingly use discretized self-supervised speech representations for compression and integration with token-based models, yet their impact on emotional inf…