5 papers
Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
Justin Lazarow, Kai Kang, Afshin Dehghan
We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric…
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik Daxberger, Nina Wenzel, David Griffiths +8
Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-q…
Cubify Anything: Scaling Indoor 3D Object Detection
Justin Lazarow, David Griffiths, Gefen Kohavi +2
We consider indoor 3D object detection with respect to a single RGB(-D) frame acquired from a commodity handheld device. We seek to significantly advance the status quo with respec…
Unaligned Image-to-Sequence Transformation with Loop Consistency
Siyang Wang, Justin Lazarow, Kwonjoon Lee +1
We tackle the problem of modeling sequential visual phenomena. Given examples of a phenomena that can be divided into discrete time steps, we aim to take an input from any such tim…
Learning Instance Occlusion for Panoptic Segmentation
Justin Lazarow, Kwonjoon Lee, Kunyu Shi +1
Panoptic segmentation requires segments of both "things" (countable object instances) and "stuff" (uncountable and amorphous regions) within a single output. A common approach invo…