7 papers
Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
Hyeonyu Kim, Sehwan Lim, Youngwon Choi +2
Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Recent visual token pruning me…
CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents
Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon +2
Task-oriented voice agents need to map spoken user requests to structured outputs such as semantic frames, executable actions, and function calls. A common approach is to cascade A…
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
Youngwon Choi, Jinwoo Oh, Hwayeon Kim +1
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide l…
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
Donghyuk Jung, Youngwon Choi
We analyze how automatic speech recognition (ASR) errors propagate through ASR--LLM cascades in Korean spoken question answering (SQA), focusing on downstream semantic failures tha…
Mine-JEPA: In-Domain Self-Supervised Learning for Mine-Like Object Classification in Side-Scan Sonar
Taeyoun Kwon, Youngwon Choi, Hyeonyu Kim +3
Side-scan sonar (SSS) mine classification is a challenging maritime vision problem characterized by extreme data scarcity and a large domain gap from natural images. While self-sup…
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
Youngwon Choi, Jaeyoon Jung, Hyeonyu Kim +2
Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge…