6 citations · 6 across the 14 of their papers we have counts for
Showing 2026 · eess.ASShow all
2 papers · 2 filters
eess.AS2026
Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over p…
eess.AS2026
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2
While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation q…