works on

From the 2 of 11 linked papers with an AI index.

activity
20242026
collaborators

11 papers

eess.AS2026

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim +8

The paper proposes using audio-aware large language models to give fine‑grained feedback on text‑to‑audio generation, improving how well the generated audio follows multi‑event and…

cs.SD2026

Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis

Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1

The paper introduces a Speech Representation Fréchet Distance loss (SR‑FD) that regularizes few‑step diffusion/flow‑matching TTS models by matching the statistics of Whisper and CT…

eess.AS2026

FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation

Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2

While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation q…

cs.SD2026

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

Kuan-Po Huang, Bo-Ru Lu, Byeonggeun Kim +8

Autoregressive (AR) models with diffusion heads have recently achieved strong text-to-audio performance, yet their iterative decoding and multi-step sampling process introduce high…

eess.AS2026

DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment

Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu +25

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following. Recent LALMs augment Large Langu…

cs.SD2026

The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

Zhixian Zhao, Shuiyuan Wang, Guojian Li +12

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowin…