collaborators

6 papers

cs.CL2026

Continuous Audio Thinking for Large Audio Language Models

Gyojin Han, Dong-Jae Lee, Changho Choi +2

Large audio language models (LALMs) have shown impressive capabilities on diverse audio understanding tasks, ranging from speech transcription to music analysis. However, because L…

cs.CV2026

Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering

Minchan Kwon, Hyounguk Shon, Junmo Kim

Large multimodal models (LMMs) have recently demonstrated remarkable performance in video question answering (VideoQA), yet reasoning over video remains challenging due to high inf…

eess.AS2025

FxSearcher: gradient-free text-driven audio transformation

Hojoon Ki, Jongsuk Kim, Minchan Kwon +1

Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set…

cs.CL2025

Preference Distillation via Value based Reinforcement Learning

Minchan Kwon, Junwon Ko, Kangil Kim +1

Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision…

eess.AS2025

FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition

Jongsuk Kim, Jaemyung Yu, Minchan Kwon +1

Large-scale ASR models have achieved remarkable gains in accuracy and robustness. However, fairness issues remain largely unaddressed despite their critical importance in real-worl…

eess.AS2025

InfiniteAudio: Infinite-Length Audio Generation with Consistency

Chaeyoung Jung, Hojoon Ki, Ji-Hoon Kim +2

This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory…