4 papers
FxSearcher: gradient-free text-driven audio transformation
Hojoon Ki, Jongsuk Kim, Minchan Kwon +1
Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set…
Preference Distillation via Value based Reinforcement Learning
Minchan Kwon, Junwon Ko, Kangil Kim +1
Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision…
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
Jongsuk Kim, Jaemyung Yu, Minchan Kwon +1
Large-scale ASR models have achieved remarkable gains in accuracy and robustness. However, fairness issues remain largely unaddressed despite their critical importance in real-worl…
InfiniteAudio: Infinite-Length Audio Generation with Consistency
Chaeyoung Jung, Hojoon Ki, Ji-Hoon Kim +2
This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory…