6 citations · 6 across the 12 of their papers we have counts for
5 papers · 1 filter
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2
While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation q…
TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
Yi-Cheng Lin, Yu-Hua Chen, Jia-Kai Dong +12
Large audio-language models are advancing rapidly, yet most evaluations emphasize speech or globally sourced sounds, overlooking culturally distinctive cues. This gap raises a crit…
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +13
Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The i…
Towards audio language modeling -- an overview
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +4
Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as su…
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
Haibin Wu, Ho-Lam Chung, Yi-Cheng Lin +7
The sound codec's dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance. Recent years have witnessed significant developme…