5 papers
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
Jelena BratuliÄ, Sudhanshu Mittal, Thomas Brox +1
Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure…
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
Jelena BratuliÄ, Sudhanshu Mittal, David T. Hoffmann +5
Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updat…
MAD-Sherlock: Multi-Agent Debate for Visual Misinformation Detection
Kumud Lakara, Georgia Channing, Christian Rupprecht +4
One of the most challenging forms of misinformation involves pairing images with misleading text to create false narratives. Existing AI-driven detection systems often require doma…
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
Jensen Zhou, Hang Gao, Vikram Voleti +6
We present Stable Virtual Camera (Seva), a generalist diffusion model that creates novel views of a scene, given any number of input views and target cameras. Existing works strugg…
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
Qianyi Deng, Oishi Deb, Amir Patel +4
Animal pose estimation (APE) aims to locate the animal body parts using a diverse array of sensor and modality inputs (e.g. RGB cameras, LiDAR, infrared, IMU, acoustic and language…