3 papers
cs.CV2026
Learned Image Compression for Vision-Language-Action Models
Hyeonjun Kim, Jegwang Ryu, Sangbeom Ha +4
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in b…
cs.RO2025
: distilling for long-horizon prehensile and non-prehensile manipulation
Haewon Jung, Donguk Lee, Haecheol Park +2
Current robots struggle with long-horizon manipulation tasks requiring sequences of prehensile and non-prehensile skills, contact-rich interactions, and long-term reasoning. We pre…
cs.CV2024
Diversify, Contextualize, and Adapt: Efficient Entropy Modeling for Neural Image Codec
Jun-Hyuk Kim, Seungeon Kim, Won-Hee Lee +1
Designing a fast and effective entropy model is challenging but essential for practical application of neural codecs. Beyond spatial autoregressive entropy models, more efficient b…