activity
20242026
collaborators

8 papers

cs.CV2026

Streaming Dense Voxel Representations for 3D Occupancy Prediction

Seokha Moon, Janghyun Baek, Yujin Jeong +5

In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-grained spatial details and streami…

cs.RO2026

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

Jiyun Jang, Yujin Sung, Woosung Joung +5

Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level ac…

cs.CV2025

SemanticControl: A Training-Free Approach for Handling Loosely Aligned Visual Conditions in ControlNet

Woosung Joung, Daewon Chae, Jinkyu Kim

ControlNet has enabled detailed spatial control in text-to-image diffusion models by incorporating additional visual conditions such as depth or edge maps. However, its effectivene…

cs.RO2025

Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent

Che Rin Yu, Daewon Chae, Dabin Seo +3

When humans perform everyday tasks, we naturally adjust our actions based on the current state of the environment. For instance, if we intend to put something into a drawer but not…

cs.LG2025

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Thomas Zeng, Shuibai Zhang, Shutong Wu +13

Process Reward Models (PRMs) have proven effective at enhancing mathematical reasoning for Large Language Models (LLMs) by leveraging increased inference-time computation. However,…

cs.CV2025

DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models

Daewon Chae, June Suk Choi, Jinkyu Kim +1

Fine-tuning text-to-image diffusion models to maximize rewards has proven effective for enhancing model performance. However, reward fine-tuning methods often suffer from slow conv…