works on

From the 2 of 13 linked papers with an AI index.

activity
20242026
most citedThe ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS2026

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim +8

The paper proposes using audio-aware large language models to give fine‑grained feedback on text‑to‑audio generation, improving how well the generated audio follows multi‑event and…

eess.AS2026

FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation

Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2

While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation q…

eess.AS2026

DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment

Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu +25

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following. Recent LALMs augment Large Langu…

eess.AS2025

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

Shu-wen Yang, Byeonggeun Kim, Kuan-Po Huang +8

Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language P…

eess.AS2025

SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition

Ming-Hao Hsu, Hung-yi Lee

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due…

eess.AS2025

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling

Kuan-Po Huang, Shu-wen Yang, Huy Phan +8

Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent t…