8 papers
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Zhengqin Li, Cheng Zhang, Jakob Engel +1
We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction. Although recent object-centric feed-forw…
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
Rosario Leonardi, Francesco Ragusa, Daniele Materia +4
Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, an…
ART: Articulated Reconstruction Transformer
Zizhang Li, Cheng Zhang, Zhengqin Li +7
We introduce ART, Articulated Reconstruction Transformer -- a category-agnostic, feed-forward model that reconstructs complete 3D articulated objects from only sparse, multi-state…
Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D
Daniel DeTone, Tianwei Shen, Fan Zhang +4
Detecting and localizing objects in space is a fundamental computer vision problem. While much progress has been made to solve 2D object detection, 3D object localization is much l…
JRM: Joint Reconstruction Model for Multiple Objects without Alignment
Qirui Wu, Yawar Siddiqui, Duncan Frost +6
Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards st…
NymeriaPlus: Enriching Nymeria Dataset with Additional Annotations and Data
Daniel DeTone, Federica Bogo, Eric-Tuan Le +7
The Nymeria Dataset, released in 2024, is a large-scale collection of in-the-wild human activities captured with multiple egocentric wearable devices that are spatially localized a…