3 papers
cs.CV2024
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
Andrew Marmon, Grant Schindler, José Lezama +3
We extend multimodal transformers to include 3D camera motion as a conditioning signal for the task of video generation. Generative video models are becoming increasingly powerful,…
cs.CV2024
SLAIM: Robust Dense Neural SLAM for Online Tracking and Mapping
Vincent Cartillier, Grant Schindler, Irfan Essa
We present SLAIM - Simultaneous Localization and Implicit Mapping. We propose a novel coarse-to-fine tracking model tailored for Neural Radiance Field SLAM (NeRF-SLAM) to achieve s…
cs.CV2024
3D Semantic MapNet: Building Maps for Multi-Object Re-Identification in 3D
Vincent Cartillier, Neha Jain, Irfan Essa
We study the task of 3D multi-object re-identification from embodied tours. Specifically, an agent is given two tours of an environment (e.g. an apartment) under two different layo…