2 citations · 4 across the 6 of their papers we have counts for
7 papers
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
Yurou Yang, Muyuan Lin, Roberto Martin-Martin +4
Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While…
Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation
Anthony Opipari, Aravindhan K Krishnan, Shreekant Gayaka +4
Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video se…
UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References
Ming-Feng Li, Xin Yang, Fu-En Wang +5
6D object pose estimation has shown strong generalizability to novel objects. However, existing methods often require either a complete, well-reconstructed 3D model or numerous ref…
Enhancing Single Image to 3D Generation using Gaussian Splatting and Hybrid Diffusion Priors
Hritam Basak, Hadi Tabatabaee, Shreekant Gayaka +6
3D object generation from a single image involves estimating the full 3D geometry and texture of unseen views from an unposed RGB image captured in the wild. Accurately reconstruct…
Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation
Anthony Opipari, Aravindhan K Krishnan, Shreekant Gayaka +4
This paper presents a method for generating large-scale datasets to improve class-agnostic video segmentation across robots with different form factors. Specifically, we consider t…
SupeRGB-D: Zero-shot Instance Segmentation in Cluttered Indoor Environments
Evin Pınar Örnek, Aravindhan K Krishnan, Shreekant Gayaka +4
Object instance segmentation is a key challenge for indoor robots navigating cluttered environments with many small objects. Limitations in 3D sensing capabilities often make it di…