3 papers
cs.CV2026
Learned Image Compression for Vision-Language-Action Models
Hyeonjun Kim, Jegwang Ryu, Sangbeom Ha +4
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in b…
cs.CV2026
A Self-Supervised Approach on Motion Calibration for Enhancing Physical Plausibility in Text-to-Motion
Gahyeon Shim, Soogeun Park, Hyemin Ahn
Generating semantically aligned human motion from textual descriptions has made rapid progress, but ensuring both semantic and physical realism in motion remains a challenge. In th…
cs.RO2025
Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement
Hogun Kee, Wooseok Oh, Minjae Kang +2
In this paper, we present the tidiness score-guided Monte Carlo tree search (TSMCTS), a novel framework designed to address the tabletop tidying up problem using only an RGB-D came…