Showing 2026Show all
3 papers · 1 filter
eess.AS2026
Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over p…
eess.AS2026
Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu +2
Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling.…
eess.AS2026
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
Wei-Ping Huang, Chee-En Yu, Guan-Ting Lin +1
Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressive models remains theoreticall…