3 papers
cs.RO2026
Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation
Chiyoung Kim, Sanghyuk Roy Choi, Minhyeok Lee
Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on rob…
cs.SD2026
Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs
Song-ha Jo, Sehyun Lee, Soyoon Kim +2
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard thi…
cs.AI2026
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
Minhyeok Lee, Chiyoung Kim, Chanhoe Gu +5
Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven…