4 papers · 1 filter
Learned Image Compression for Vision-Language-Action Models
Hyeonjun Kim, Jegwang Ryu, Sangbeom Ha +4
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in b…
A Self-Supervised Approach on Motion Calibration for Enhancing Physical Plausibility in Text-to-Motion
Gahyeon Shim, Soogeun Park, Hyemin Ahn
Generating semantically aligned human motion from textual descriptions has made rapid progress, but ensuring both semantic and physical realism in motion remains a challenge. In th…
A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis
Esteve Valls Mascaro, Hyemin Ahn, Dongheui Lee
The synthesis of human motion has traditionally been addressed through task-dependent models that focus on specific challenges, such as predicting future motions or filling in inte…
Human-Object Interaction Prediction in Videos through Gaze Following
Zhifan Ni, Esteve Valls Mascaró, Hyemin Ahn +1
Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from ima…