5 papers
StreamAlign: Streaming Text-Aligned Speech Tokenization
Kang-wook Kim, Jinyoung Park, Jinsoo Kim +3
Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rel…
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
Takyoung Kim, Kang-wook Kim, Sang Hoon Woo +3
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardle…
SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
Keonhee Park, Gunhee Kim
Multimodal Continual Instruction Tuning (MCIT) is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving a sequence of downstream tasks. Prior methods mostly uti…
Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech
Sang Hoon Woo, Sehun Lee, Kang-wook Kim +1
Spoken dialogue systems increasingly employ large language models (LLMs) to leverage their advanced reasoning capabilities. However, direct application of LLMs in spoken communicat…
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
Wonkwang Lee, Jongwon Jeong, Taehong Moon +4
Motion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive moti…