6 papers
Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
Hyeonyu Kim, Sehwan Lim, Youngwon Choi +2
Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Recent visual token pruning me…
CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents
Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon +2
Task-oriented voice agents need to map spoken user requests to structured outputs such as semantic frames, executable actions, and function calls. A common approach is to cascade A…
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
Youngwon Choi, Jinwoo Oh, Hwayeon Kim +1
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide l…
Mine-JEPA: In-Domain Self-Supervised Learning for Mine-Like Object Classification in Side-Scan Sonar
Taeyoun Kwon, Youngwon Choi, Hyeonyu Kim +3
Side-scan sonar (SSS) mine classification is a challenging maritime vision problem characterized by extreme data scarcity and a large domain gap from natural images. While self-sup…
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
Youngwon Choi, Jaeyoon Jung, Hyeonyu Kim +2
Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge…
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
Hyeonyu Kim, Seokhoon Jeong, Seonghee Han +2
Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlightin…