From the 2 of 6 linked papers with an AI index.
6 papers
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim +8
The paper proposes using audio-aware large language models to give fine‑grained feedback on text‑to‑audio generation, improving how well the generated audio follows multi‑event and…
Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1
The paper introduces a Speech Representation Fréchet Distance loss (SR‑FD) that regularizes few‑step diffusion/flow‑matching TTS models by matching the statistics of Whisper and CT…
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2
While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation q…
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation
Kuan-Po Huang, Bo-Ru Lu, Byeonggeun Kim +8
Autoregressive (AR) models with diffusion heads have recently achieved strong text-to-audio performance, yet their iterative decoding and multi-step sampling process introduce high…
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
Shu-wen Yang, Byeonggeun Kim, Kuan-Po Huang +8
Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language P…
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
Kuan-Po Huang, Shu-wen Yang, Huy Phan +8
Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent t…