56 papers
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
Donghu Kim, Youngdo Lee, Hojoon Lee +6
Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challeng…
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
Jooyeol Yun, Jintae Park, Hyesu Lim +3
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering…
Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
Daehoon Gwak, Minhyung Lee, Junwoo Park +1
The paper surveys methods for speeding up inference of masked diffusion large language models by categorizing algorithmic, architectural, and system-level acceleration techniques a…
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Byungkun Lee, Dongyoon Hwang, Dongjin Kim +3
The paper proposes robot-centric pointmaps, which encode 3D scene coordinates in the robot's frame as image pixels, enabling vision‑language‑action models to align visual inputs wi…
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Dongyoon Hwang, Byungkun Lee, Dongjin Kim +7
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm u…
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Hoiyeong Jin, Hyojin Jang, Junha Hyung +6
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D s…