11 citations · 55 across the 50 of their papers we have counts for
9 papers · 2 filters
Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over p…
AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning
Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang +2
In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition…
Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu +2
Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling.…
Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach
Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou +4
As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV…
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
Yi-Cheng Lin, Yun-Shao Tsai, Kuan-Yu Chen +6
Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-lear…
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou +5
Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To q…