papers

Publications (162)

cs.CV2025

Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

Yuanbo Yang, Jiahao Shao, Xinyang Li +3

In this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as m…

cs.CV2014

FollowMe: Efficient Online Min-Cost Flow Tracking with Bounded Memory and Computation

Philip Lenz, Andreas Geiger, Raquel Urtasun

One of the most popular approaches to multi-target tracking is tracking-by-detection. Current min-cost flow algorithms which solve the data association problem optimally have three…

cs.CV2026

FrankenMotion: Part-level Human Motion Generation and Composition

Chuqiao Li, Xianghui Xie, Yong Cao +2

Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptio…

cs.CV2025

Level-Set Parameters: Novel Representation for 3D Shape Analysis

Huan Lei, Hongdong Li, Andreas Geiger +1

3D shape analysis has been largely focused on traditional 3D representations of point clouds and meshes, but the discrete nature of these data makes the analysis susceptible to var…

cs.CV2022

TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving

Kashyap Chitta, Aditya Prakash, Bernhard Jaeger +3

How should we integrate representations from complementary sensors for autonomous driving? Geometry-based fusion has shown promise for perception (e.g. object detection, motion for…

cs.CV2021

Locally Aware Piecewise Transformation Fields for 3D Human Mesh Registration

Shaofei Wang, Andreas Geiger, Siyu Tang

Registering point clouds of dressed humans to parametric human models is a challenging task in computer vision. Traditional approaches often rely on heavily engineered pipelines th…

cs.LG2018

Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks

Lars Mescheder, Sebastian Nowozin, Andreas Geiger

Variational Autoencoders (VAEs) are expressive latent variable models that can be used to learn complex probability distributions from training data. However, the quality of the re…

cs.CV2022

Fast-SNARF: A Fast Deformer for Articulated Neural Fields

Xu Chen, Tianjian Jiang, Jie Song +4

Neural fields have revolutionized the area of 3D reconstruction and novel view synthesis of rigid scenes. A key challenge in making such methods applicable to articulated objects,…

cs.CV2018

Semantic Visual Localization

Johannes L. Schönberger, Marc Pollefeys, Andreas Geiger +1

Robust visual localization under a wide range of viewing conditions is a fundamental problem in computer vision. Handling the difficult cases of this problem is not only very chall…

cs.LG2026

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation

Jonathan von Rad, Yong Cao, Andreas Geiger

Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primari…

cs.CV2022

ARAH: Animatable Volume Rendering of Articulated Human SDFs

Shaofei Wang, Katja Schwarz, Andreas Geiger +1

Combining human body models with differentiable rendering has recently enabled animatable avatars of clothed humans from sparse sets of multi-view RGB videos. While state-of-the-ar…

cs.CV2024

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

Yuedong Chen, Haofei Xu, Chuanxia Zheng +5

We introduce MVSplat, an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we bu…

cs.LG2022

StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets

Axel Sauer, Katja Schwarz, Andreas Geiger

Computer graphics has experienced a recent surge of data-centric approaches for photorealistic and controllable content creation. StyleGAN in particular sets new standards for gene…

cs.RO2018

Conditional Affordance Learning for Driving in Urban Environments

Axel Sauer, Nikolay Savinov, Andreas Geiger

Most existing approaches to autonomous driving fall into one of two categories: modular pipelines, that build an extensive model of the environment, and imitation learning approach…

cs.CV2023

On Offline Evaluation of 3D Object Detection for Autonomous Driving

Tim Schreier, Katrin Renz, Andreas Geiger +1

Prior work in 3D object detection evaluates models using offline metrics like average precision since closed-loop online evaluation on the downstream driving task is costly. Howeve…

cs.CV2023

Deep Generative Models on 3D Representations: A Survey

Zifan Shi, Sida Peng, Yinghao Xu +3

Generative models aim to learn the distribution of observed data by generating new instances. With the advent of neural networks, deep generative models, including variational auto…

cs.CV2019

Superquadrics Revisited: Learning 3D Shape Parsing beyond Cuboids

Despoina Paschalidou, Ali Osman Ulusoy, Andreas Geiger

Abstracting complex 3D shapes with parsimonious part-based representations has been a long standing goal in computer vision. This paper presents a learning-based solution to this p…

cs.CV2024

GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers

Takeru Miyato, Bernhard Jaeger, Max Welling +1

As transformers are equivariant to the permutation of input tokens, encoding the positional information of tokens is necessary for many tasks. However, since existing positional en…

cs.RO2025

PlanT 2.0: Exposing Biases and Structural Flaws in Closed-Loop Driving

Simon Gerstenecker, Andreas Geiger, Katrin Renz

Most recent work in autonomous driving has prioritized benchmark performance and methodological innovation over in-depth analysis of model failures, biases, and shortcut learning.…

cs.CV2025

Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation

Volodymyr Havrylov, Haiwen Huang, Dan Zhang +1

Vision Foundation Models (VFMs) are large-scale, pre-trained models that serve as general-purpose backbones for various computer vision tasks. As VFMs' popularity grows, there is a…

cs.LG2023

StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

Axel Sauer, Tero Karras, Samuli Laine +2

Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model familie…

cs.CV2026

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

Marcel Plocher, Bernhard Schölkopf, Andreas Geiger +1

The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumption is particularly questiona…

cs.RO2026

Fail2Drive: Benchmarking Closed-Loop Driving Generalization

Simon Gerstenecker, Andreas Geiger, Katrin Renz

Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing b…

cs.CV2021

GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields

Michael Niemeyer, Andreas Geiger

Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable.…

cs.LG2025

CaRL: Learning Scalable Planning Policies with Simple Rewards

Bernhard Jaeger, Daniel Dauner, Jens Beißwenger +3

We investigate reinforcement learning (RL) for privileged planning in autonomous driving. State-of-the-art approaches for this task are rule-based, but these methods do not scale t…

cs.CV2024

NeLF-Pro: Neural Light Field Probes for Multi-Scale Novel View Synthesis

Zinuo You, Andreas Geiger, Anpei Chen

We present NeLF-Pro, a novel representation to model and reconstruct light fields in diverse natural scenes that vary in extent and spatial granularity. In contrast to previous fas…

cs.CV2021

KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs

Christian Reiser, Songyou Peng, Yiyi Liao +1

NeRF synthesizes novel views of a scene with unprecedented quality by fitting a neural radiance field to RGB images. However, NeRF requires querying a deep Multi-Layer Perceptron (…

cs.CV2024

Unimotion: Unifying 3D Human Motion Synthesis and Understanding

Chuqiao Li, Julian Chibane, Yannan He +3

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control a…

cs.CV2023

MERF: Memory-Efficient Radiance Fields for Real-time View Synthesis in Unbounded Scenes

Christian Reiser, Richard Szeliski, Dor Verbin +5

Neural radiance fields enable state-of-the-art photorealistic view synthesis. However, existing radiance field representations are either too compute-intensive for real-time render…

cs.CV2022

TensoRF: Tensorial Radiance Fields

Anpei Chen, Zexiang Xu, Andreas Geiger +2

We present TensoRF, a novel approach to model and reconstruct radiance fields. Unlike NeRF that purely uses MLPs, we model the radiance field of a scene as a 4D tensor, which repre…

cs.CV2021

CAMPARI: Camera-Aware Decomposed Generative Neural Radiance Fields

Michael Niemeyer, Andreas Geiger

Tremendous progress in deep generative models has led to photorealistic image synthesis. While achieving compelling results, most approaches operate in the two-dimensional image do…

cs.CV2025

Efficient Depth-Guided Urban View Synthesis

Sheng Miao, Jiaxin Huang, Dongfeng Bai +4

Recent advances in implicit scene representation enable high-fidelity street view novel view synthesis. However, existing methods optimize a neural radiance field for each scene, r…

cs.CV2022

VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel Grids

Katja Schwarz, Axel Sauer, Michael Niemeyer +2

State-of-the-art 3D-aware generative models rely on coordinate-based MLPs to parameterize 3D radiance fields. While demonstrating impressive results, querying an MLP for every samp…

cs.CV2024

NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

Daniel Dauner, Marcel Hallgarten, Tianyu Li +9

Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On th…

cs.CV2023

GOOD: Exploring Geometric Cues for Detecting Objects in an Open World

Haiwen Huang, Andreas Geiger, Dan Zhang

We address the task of open-world class-agnostic object detection, i.e., detecting every object in an image by learning from a limited number of base object classes. State-of-the-a…

cs.CV2021

SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes

Xu Chen, Yufeng Zheng, Michael J. Black +2

Neural implicit surface representations have emerged as a promising paradigm to capture 3D shapes in a continuous and resolution-independent manner. However, adapting them to artic…

cs.CV2021

ATISS: Autoregressive Transformers for Indoor Scene Synthesis

Despoina Paschalidou, Amlan Kar, Maria Shugrina +3

The ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to dat…

cs.CV2024

LaRa: Efficient Large-Baseline Radiance Fields

Anpei Chen, Haofei Xu, Stefano Esposito +2

Radiance field methods have achieved photorealistic novel view synthesis and geometry reconstruction. But they are mostly applied in per-scene optimization or small-baseline settin…

cs.CV2020

Category Level Object Pose Estimation via Neural Analysis-by-Synthesis

Xu Chen, Zijian Dong, Jie Song +2

Many object pose estimation algorithms rely on the analysis-by-synthesis framework which requires explicit representations of individual object instances. In this paper we combine…

cs.CV2020

Learning Unsupervised Hierarchical Part Decomposition of 3D Objects from a Single RGB Image

Despoina Paschalidou, Luc van Gool, Andreas Geiger

Humans perceive the 3D world as a set of distinct objects that are characterized by various low-level (geometry, reflectance) and high-level (connectivity, adjacency, symmetry) pro…

cs.RO2025

Is Single-View Mesh Reconstruction Ready for Robotics?

Frederik Nolte, Andreas Geiger, Bernhard Schölkopf +1

This paper evaluates single-view mesh reconstruction models for their potential in enabling instant digital twin creation for real-time planning and dynamics prediction using physi…

cs.CV2024

Renovating Names in Open-Vocabulary Segmentation Benchmarks

Haiwen Huang, Songyou Peng, Dan Zhang +1

Names are essential to both human cognition and vision-language models. Open-vocabulary models utilize class names as text prompts to generalize to categories unseen during trainin…

cs.CV2021

Learning Cascaded Detection Tasks with Weakly-Supervised Domain Adaptation

Niklas Hanselmann, Nick Schneider, Benedikt Ortelt +1

In order to handle the challenges of autonomous driving, deep learning has proven to be crucial in tackling increasingly complex tasks, such as 3D detection or instance segmentatio…

cs.CV2025

Relightable Full-Body Gaussian Codec Avatars

Shaofei Wang, Tomas Simon, Igor Santesteban +15

We propose Relightable Full-Body Gaussian Codec Avatars, a new approach for modeling relightable full-body avatars with fine-grained details including face and hands. The unique ch…

cs.CV2020

Towards Unsupervised Learning of Generative Models for 3D Controllable Image Synthesis

Yiyi Liao, Katja Schwarz, Lars Mescheder +1

In recent years, Generative Adversarial Networks have achieved impressive results in photorealistic image synthesis. This progress nurtures hopes that one day the classical renderi…

cs.CV2025

2D Gaussian Splatting for Geometrically Accurate Radiance Fields

Binbin Huang, Zehao Yu, Anpei Chen +2

3D Gaussian Splatting (3DGS) has recently revolutionized radiance field reconstruction, achieving high quality novel view synthesis and fast rendering speed without baking. However…

cs.CV2024

HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting

Hongyu Zhou, Jiahao Shao, Lu Xu +6

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel vie…

cs.RO2026

123D: Unifying Multi-Modal Autonomous Driving Data at Scale

Daniel Dauner, Valentin Charraut, Bastian Berle +10

The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain largely untapped. Each dataset…

cs.CV2023

AG3D: Learning to Generate 3D Avatars from 2D Image Collections

Zijian Dong, Xu Chen, Jinlong Yang +3

While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing meth…

cs.CV2026

Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

Weijie Wang, Qihang Cao, Sensen Gao +10

Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and interacting with the physical…

cs.CV2025

Easi3R: Estimating Disentangled Motion from DUSt3R Without Training

Xingyu Chen, Yue Chen, Yuliang Xiu +2

Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct super…

cs.LG2018

Which Training Methods for GANs do actually Converge?

Lars Mescheder, Andreas Geiger, Sebastian Nowozin

Recent work has shown local convergence of GAN training for absolutely continuous data and generator distributions. In this paper, we show that the requirement of absolute continui…

cs.CV2024

Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes

Zehao Yu, Torsten Sattler, Andreas Geiger

Recently, 3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis results, while allowing the rendering of high-resolution images in real-time. However, lever…

cs.CV2024

MuRF: Multi-Baseline Radiance Fields

Haofei Xu, Anpei Chen, Yuedong Chen +5

We present Multi-Baseline Radiance Fields (MuRF), a general feed-forward approach to solving sparse view synthesis under multiple different baseline settings (small and large basel…

cs.CV2023

Mip-Splatting: Alias-free 3D Gaussian Splatting

Zehao Yu, Anpei Chen, Binbin Huang +2

Recently, 3D Gaussian Splatting has demonstrated impressive novel view synthesis results, reaching high fidelity and efficiency. However, strong artifacts can be observed when chan…

cs.CV2021

SMD-Nets: Stereo Mixture Density Networks

Fabio Tosi, Yiyi Liao, Carolin Schmitt +1

Despite stereo matching accuracy has greatly improved by deep learning in the last few years, recovering sharp boundaries and high-resolution outputs efficiently remains challengin…

cs.CV2026

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

Peizheng Li, Zhenghao Zhang, David Holtz +6

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning ca…

cs.CV2026

LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving

Long Nguyen, Micha Fauth, Bernhard Jaeger +4

Simulators can generate virtually unlimited driving data, yet imitation learning policies in simulation still struggle to achieve robust closed-loop performance. Motivated by this…

cs.CV2024

GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs

Gege Gao, Weiyang Liu, Anpei Chen +2

As pretrained text-to-image diffusion models become increasingly powerful, recent efforts have been made to distill knowledge from these text-to-image pretrained models for optimiz…

cs.CV2022

PINA: Learning a Personalized Implicit Neural Avatar from a Single RGB-D Video Sequence

Zijian Dong, Chen Guo, Jie Song +3

We present a novel method to learn Personalized Implicit Neural Avatars (PINA) from a short RGB-D sequence. This allows non-expert users to create a detailed and personalized virtu…

cs.CV2021

Shape As Points: A Differentiable Poisson Solver

Songyou Peng, Chiyu "Max" Jiang, Yiyi Liao +3

In recent years, neural implicit representations gained popularity in 3D reconstruction due to their expressiveness and flexibility. However, the implicit nature of neural implicit…

cs.CV2023

Towards Scalable Multi-View Reconstruction of Geometry and Materials

Carolin Schmitt, Božidar Antić, Andrei Neculai +2

In this paper, we propose a novel method for joint recovery of camera pose, object geometry and spatially-varying Bidirectional Reflectance Distribution Function (svBRDF) of 3D sce…

cs.RO2025

Centaur: Robust End-to-End Autonomous Driving with Test-Time Training

Chonghao Sima, Kashyap Chitta, Zhiding Yu +5

How can we rely on an end-to-end autonomous vehicle's complex decision-making system during deployment? One common solution is to have a ``fallback layer'' that checks the planned…

cs.CV2018

Learning 3D Shape Completion under Weak Supervision

David Stutz, Andreas Geiger

We address the problem of 3D shape completion from sparse and noisy point clouds, a fundamental problem in computer vision and robotics. Recent approaches are either data-driven or…

cs.CV2017

Augmented Reality Meets Computer Vision : Efficient Data Generation for Urban Driving Scenes

Hassan Abu Alhaija, Siva Karthik Mustikovela, Lars Mescheder +2

The success of deep learning in computer vision is based on availability of large annotated datasets. To lower the need for hand labeled images, virtually rendered 3D worlds have r…

cs.CV2026

EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis

Sheng Miao, Sijin Li, Pan Wang +5

Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time wit…

cs.LG2025

An Invitation to Deep Reinforcement Learning

Bernhard Jaeger, Andreas Geiger

Training a deep neural network to maximize a target objective has become the standard recipe for successful machine learning over the last decade. These networks can be optimized w…

cs.LG2024

HDT: Hierarchical Document Transformer

Haoyu He, Markus Flicke, Jan Buchmann +2

In this paper, we propose the Hierarchical Document Transformer (HDT), a novel sparse Transformer architecture tailored for structured hierarchical documents. Such documents are ex…

cs.CV2020

Convolutional Occupancy Networks

Songyou Peng, Michael Niemeyer, Lars Mescheder +2

Recently, implicit neural representations have gained popularity for learning-based 3D reconstruction. While demonstrating promising results, most implicit approaches are limited t…

cs.CV2020

Learning Implicit Surface Light Fields

Michael Oechsle, Michael Niemeyer, Lars Mescheder +2

Implicit representations of 3D objects have recently achieved impressive results on learning-based 3D reconstruction tasks. While existing works use simple texture models to repres…

cs.CV2026

InvSplat: Inverse Feed-Forward Scene Splatting

Polina Karpikova, Wenjing Bian, Haofei Xu +2

Inverse rendering aims to recover both 3D geometry and physically meaningful material properties from images, enabling applications such as relighting and novel view synthesis. Opt…

cs.CV2021

Benchmarking Unsupervised Object Representations for Video Sequences

Marissa A. Weis, Kashyap Chitta, Yash Sharma +4

Perceiving the world in terms of objects and tracking them through time is a crucial prerequisite for reasoning and scene understanding. Recently, several methods have been propose…

cs.CV2021

Intrinsic Autoencoders for Joint Neural Rendering and Intrinsic Image Decomposition

Hassan Abu Alhaija, Siva Karthik Mustikovela, Justus Thies +4

Neural rendering techniques promise efficient photo-realistic image synthesis while at the same time providing rich control over scene parameters by learning the physical image for…

cs.CV2025

DepthSplat: Connecting Gaussian Splatting and Depth

Haofei Xu, Songyou Peng, Fangjinhua Wang +4

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and…

cs.CV2026

Slots, Transitions, Loops: Learning Composable World Models for ARC

Gege Gao, Bernhard Schölkopf, Andreas Geiger

ARC tests in-context rule induction: given a few input-output demonstrations, a model must infer the hidden rule and apply it to a new query. While many approaches express ARC rule…

cs.CV2021

GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis

Katja Schwarz, Yiyi Liao, Michael Niemeyer +1

While 2D generative adversarial networks have enabled high-resolution image synthesis, they largely lack an understanding of the 3D world and the image formation process. Thus, the…

cs.CV2020

Self-Supervised Linear Motion Deblurring

Peidong Liu, Joel Janai, Marc Pollefeys +2

Motion blurry images challenge many computer vision algorithms, e.g, feature detection, motion estimation, or object recognition. Deep convolutional neural networks are state-of-th…

cs.CV2025

Volumetric Surfaces: Representing Fuzzy Geometries with Layered Meshes

Stefano Esposito, Anpei Chen, Christian Reiser +7

High-quality view synthesis relies on volume rendering, splatting, or surface rendering. While surface rendering is typically the fastest, it struggles to accurately model fuzzy ge…

cs.LG2025

Artificial Kuramoto Oscillatory Neurons

Takeru Miyato, Sindy Löwe, Andreas Geiger +1

It has long been known in both neuroscience and AI that ``binding'' between neurons leads to a form of competitive learning where representations are compressed in order to represe…

cs.CV2019

RayNet: Learning Volumetric 3D Reconstruction with Ray Potentials

Despoina Paschalidou, Ali Osman Ulusoy, Carolin Schmitt +2

In this paper, we consider the problem of reconstructing a dense 3D model using images captured from different views. Recent methods based on convolutional neural networks (CNN) al…

cs.CV2026

Learn2Splat: Extending the Horizon of Learned 3DGS Optimization

Naama Pearl, Stefano Esposito, Haofei Xu +6

3D Gaussian Splatting (3DGS) optimization is most commonly performed using standard optimizers (Adam, SGD). While stable across diverse scenes, standard optimizers are general-purp…

cs.CV2025

sshELF: Single-Shot Hierarchical Extrapolation of Latent Features for 3D Reconstruction from Sparse-Views

Eyvaz Najafli, Marius Kästingschäfer, Sebastian Bernhard +2

Reconstructing unbounded outdoor scenes from sparse outward-facing views poses significant challenges due to minimal view overlap. Previous methods often lack cross-scene understan…

cs.CV2022

KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D

Yiyi Liao, Jun Xie, Andreas Geiger

For the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other…

cs.CV2019

MOTS: Multi-Object Tracking and Segmentation

Paul Voigtlaender, Michael Krause, Aljosa Osep +4

This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two ex…

cs.CV2023

Unifying Flow, Stereo and Depth Estimation

Haofei Xu, Jing Zhang, Jianfei Cai +4

We present a unified formulation and model for three motion and 3D perception tasks: optical flow, rectified stereo matching and unrectified stereo depth estimation from posed imag…

cs.CV2026

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

Merve Kocabas, Gege Gao, Bernhard Schölkopf +1

Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, wh…

cs.CV2021

RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs

Michael Niemeyer, Jonathan T. Barron, Ben Mildenhall +3

Neural Radiance Fields (NeRF) have emerged as a powerful representation for the task of novel view synthesis due to their simplicity and state-of-the-art performance. Though NeRF c…

cs.RO2022

PlanT: Explainable Planning Transformers via Object-Level Representations

Katrin Renz, Kashyap Chitta, Otniel-Bogdan Mercea +3

Planning an optimal route in a complex environment requires efficient reasoning about the surrounding scene. While human drivers prioritize important objects and ignore details not…

cs.CV2024

HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving

Hongyu Zhou, Longzhong Lin, Jiabao Wang +6

In the past few decades, autonomous driving algorithms have made significant progress in perception, planning, and control. However, evaluating individual components does not fully…

cs.CV2025

MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction

Zijian Dong, Longteng Duan, Jie Song +2

We present MoGA, a novel method to reconstruct high-fidelity 3D Gaussian avatars from a single-view image. The main challenge lies in inferring unseen appearance and geometric deta…

cs.CV2020

Learning Neural Light Transport

Paul Sanzenbacher, Lars Mescheder, Andreas Geiger

In recent years, deep generative models have gained significance due to their ability to synthesize natural-looking images with applications ranging from virtual reality to data au…

cs.CV2024

Hidden Biases of End-to-End Driving Datasets

Julian Zimmerlin, Jens Beißwenger, Bernhard Jaeger +2

End-to-end driving systems have made rapid progress, but have so far not been applied to the challenging new CARLA Leaderboard 2.0. Further, while there is a large body of literatu…

cs.CV2020

Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision

Michael Niemeyer, Lars Mescheder, Michael Oechsle +1

Learning-based 3D reconstruction methods have shown impressive results. However, most methods require 3D supervision which is often hard to obtain for real-world datasets. Recently…

cs.RO2026

World Engine: Towards the Era of Post-Training for Autonomous Driving

Tianyu Li, Li Chen, Caojun Wang +16

Autonomous vehicles must operate safely in the real world, where errors can have severe consequences. Although modern end-to-end driving policies excel in routine scenarios, their…

cs.CV2017

On the Integration of Optical Flow and Action Recognition

Laura Sevilla-Lara, Yiyi Liao, Fatma Guney +3

Most of the top performing action recognition methods use optical flow as a "black box" input. Here we take a deeper look at the combination of flow and action recognition, and inv…

cs.CV2024

3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting

Zhiyin Qian, Shaofei Wang, Marko Mihajlovic +2

We introduce an approach that creates animatable human avatars from monocular videos using 3D Gaussian Splatting (3DGS). Existing methods based on neural radiance fields (NeRFs) ac…

cs.CV2020

Label Efficient Visual Abstractions for Autonomous Driving

Aseem Behl, Kashyap Chitta, Aditya Prakash +2

It is well known that semantic segmentation can be used as an effective intermediate representation for learning driving policies. However, the task of street scene semantic segmen…

cs.CV2021

Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

Aditya Prakash, Kashyap Chitta, Andreas Geiger

How should representations from complementary sensors be integrated for autonomous driving? Geometry-based sensor fusion has shown great promise for perception tasks such as object…

cs.CV2022

gDNA: Towards Generative Detailed Neural Avatars

Xu Chen, Tianjian Jiang, Jie Song +4

To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging…

cs.LG2021

Counterfactual Generative Networks

Axel Sauer, Andreas Geiger

Neural networks are prone to learning shortcuts -- they often model simple correlations, ignoring more complex ones that potentially generalize better. Prior works on image classif…