9 papers · 1 filter
SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
Siyi Chen, Mikaela Angelina Uy, Chan Hee Song +6
Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The ag…
Learning Convex Decomposition via Feature Fields
Yuezhi Yang, Qixing Huang, Mikaela Angelina Uy +1
This work proposes a new formulation to the long-standing problem of convex decomposition through learning feature fields, enabling the first feed-forward model for open-world conv…
Img2CAD: Reverse Engineering 3D CAD Models from Images through VLM-Assisted Conditional Factorization
Yang You, Mikaela Angelina Uy, Jiaqi Han +7
Reverse engineering 3D computer-aided design (CAD) models from images is an important task for many downstream applications including interactive editing, manufacturing, architectu…
Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
Yiqing Liang, Mikhail Okunev, Mikaela Angelina Uy +4
Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is int…
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
Phillip Y. Lee, Jihyeon Je, Chanho Park +3
We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environmen…
PARTFIELD: Learning 3D Feature Fields for Part Segmentation and Beyond
Minghua Liu, Mikaela Angelina Uy, Donglai Xiang +4
We propose PartField, a feedforward approach for learning part-based 3D features, which captures the general concept of parts and their hierarchy without relying on predefined temp…