11 papers
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
Chan Yeong Hwang, Miso Choi, Sunghyun On +2
Understanding visual scenes requires not only recognizing objects but also reasoning about their spatial relationships. Unlike general vision-language tasks, spatial reasoning requ…
Streaming Dense Voxel Representations for 3D Occupancy Prediction
Seokha Moon, Janghyun Baek, Yujin Jeong +5
In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-grained spatial details and streami…
The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages
Miso Choi, Seonga Choi, Mincheol Kwon +3
Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, forming distinct model lineages. It rema…
AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation
Jiyun Jang, Yujin Sung, Woosung Joung +5
Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level ac…
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Kanghyun Baek, Jaihyun Lew, Chaehun Shin +2
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects…
Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling
Seokha Moon, Minseung Lee, Joon Seo +2
End-to-end autonomous driving, which bypasses traditional modular pipelines by directly predicting future trajectories from sensor inputs, has recently achieved substantial progres…