activity
20242026
most citedTowards audio language modeling -- an overview

6 citations · 7 across the 29 of their papers we have counts for

collaborators
Showing eess.ASShow all

30 papers · 1 filter

eess.AS2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang +2

In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition…

eess.AS2026

Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models

Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu +2

Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling.…

eess.AS2026

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou +4

As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV…

eess.AS2026

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI

Yi-Cheng Lin, Yun-Shao Tsai, Kuan-Yu Chen +6

Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-lear…

eess.AS2026

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou +5

Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To q…

eess.AS2026

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang +1

Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely…