Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
FxSearcher: gradient-free text-driven audio transformation
Hojoon Ki, Jongsuk Kim, Minchan Kwon +1
Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set…
eess.AS2025
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
Jongsuk Kim, Jaemyung Yu, Minchan Kwon +1
Large-scale ASR models have achieved remarkable gains in accuracy and robustness. However, fairness issues remain largely unaddressed despite their critical importance in real-worl…
eess.AS2024
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
Jongsuk Kim, Jiwon Shin, Junmo Kim
In recent years, advancements in representation learning and language models have propelled Automated Captioning (AC) to new heights, enabling the generation of human-level descrip…