20 papers
HOPE: Hand-Object Pressure Estimation from Monocular Videos
Subin Jeon, Byungjun Kim, Hanbyul Joo
Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely li…
Robot-Factored World Models via Robot Rendering
Byungjun Kim, Taeksoo Kim, Hyunsoo Cha +1
Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence future observations through two d…
Learning to Generate Human-Human-Object Interactions from Textual Descriptions
Jeonghyeon Na, Sangwon Baik, Inhee Lee +2
The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines…
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
Sangwon Baik, Gunhee Kim, Mingi Choi +1
Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fail to infer a text-consistent…
OmniRobotHome: A Multi-Camera Home Platform for Real-Time Human-Robot Interaction
Junyoung Lee, Inhee Lee, Sookwan Han +9
Robots in homes must continuously sense the people around them, yet most prior work relies on limited or offline perception. We argue that perception quality is the dominant factor…
AutoDex: An Automated Real-World System for Dexterous Grasping Data Collection
Mingi Choi, Gunhee Kim, Jisoo Kim +4
Learning robust dexterous grasping requires real-world data that records the physical outcomes of grasp attempts. Such data is hard to obtain at scale: teleoperation yields valid p…