Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Learning to Generate Human-Human-Object Interactions from Textual Descriptions
Jeonghyeon Na, Sangwon Baik, Inhee Lee +2
The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines…
cs.CV2026
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
Sangwon Baik, Gunhee Kim, Mingi Choi +1
Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fail to infer a text-consistent…
cs.CV2026
SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image
Inhee Lee, Sangwon Baik, Sungjoo Kim +3
Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation. While recent single-image lifters recover plausible pe…