Publications (162)
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
Yuanbo Yang, Jiahao Shao, Xinyang Li +3
In this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as m…
FollowMe: Efficient Online Min-Cost Flow Tracking with Bounded Memory and Computation
Philip Lenz, Andreas Geiger, Raquel Urtasun
One of the most popular approaches to multi-target tracking is tracking-by-detection. Current min-cost flow algorithms which solve the data association problem optimally have three…
FrankenMotion: Part-level Human Motion Generation and Composition
Chuqiao Li, Xianghui Xie, Yong Cao +2
Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptio…
Level-Set Parameters: Novel Representation for 3D Shape Analysis
Huan Lei, Hongdong Li, Andreas Geiger +1
3D shape analysis has been largely focused on traditional 3D representations of point clouds and meshes, but the discrete nature of these data makes the analysis susceptible to var…
TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger +3
How should we integrate representations from complementary sensors for autonomous driving? Geometry-based fusion has shown promise for perception (e.g. object detection, motion for…
Locally Aware Piecewise Transformation Fields for 3D Human Mesh Registration
Shaofei Wang, Andreas Geiger, Siyu Tang
Registering point clouds of dressed humans to parametric human models is a challenging task in computer vision. Traditional approaches often rely on heavily engineered pipelines th…
Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks
Lars Mescheder, Sebastian Nowozin, Andreas Geiger
Variational Autoencoders (VAEs) are expressive latent variable models that can be used to learn complex probability distributions from training data. However, the quality of the re…
Fast-SNARF: A Fast Deformer for Articulated Neural Fields
Xu Chen, Tianjian Jiang, Jie Song +4
Neural fields have revolutionized the area of 3D reconstruction and novel view synthesis of rigid scenes. A key challenge in making such methods applicable to articulated objects,…
Semantic Visual Localization
Johannes L. Schönberger, Marc Pollefeys, Andreas Geiger +1
Robust visual localization under a wide range of viewing conditions is a fundamental problem in computer vision. Handling the difficult cases of this problem is not only very chall…
UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
Jonathan von Rad, Yong Cao, Andreas Geiger
Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primari…
ARAH: Animatable Volume Rendering of Articulated Human SDFs
Shaofei Wang, Katja Schwarz, Andreas Geiger +1
Combining human body models with differentiable rendering has recently enabled animatable avatars of clothed humans from sparse sets of multi-view RGB videos. While state-of-the-ar…
MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
Yuedong Chen, Haofei Xu, Chuanxia Zheng +5
We introduce MVSplat, an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we bu…
StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets
Axel Sauer, Katja Schwarz, Andreas Geiger
Computer graphics has experienced a recent surge of data-centric approaches for photorealistic and controllable content creation. StyleGAN in particular sets new standards for gene…
Conditional Affordance Learning for Driving in Urban Environments
Axel Sauer, Nikolay Savinov, Andreas Geiger
Most existing approaches to autonomous driving fall into one of two categories: modular pipelines, that build an extensive model of the environment, and imitation learning approach…
On Offline Evaluation of 3D Object Detection for Autonomous Driving
Tim Schreier, Katrin Renz, Andreas Geiger +1
Prior work in 3D object detection evaluates models using offline metrics like average precision since closed-loop online evaluation on the downstream driving task is costly. Howeve…
Deep Generative Models on 3D Representations: A Survey
Zifan Shi, Sida Peng, Yinghao Xu +3
Generative models aim to learn the distribution of observed data by generating new instances. With the advent of neural networks, deep generative models, including variational auto…
Superquadrics Revisited: Learning 3D Shape Parsing beyond Cuboids
Despoina Paschalidou, Ali Osman Ulusoy, Andreas Geiger
Abstracting complex 3D shapes with parsimonious part-based representations has been a long standing goal in computer vision. This paper presents a learning-based solution to this p…
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
Takeru Miyato, Bernhard Jaeger, Max Welling +1
As transformers are equivariant to the permutation of input tokens, encoding the positional information of tokens is necessary for many tasks. However, since existing positional en…
PlanT 2.0: Exposing Biases and Structural Flaws in Closed-Loop Driving
Simon Gerstenecker, Andreas Geiger, Katrin Renz
Most recent work in autonomous driving has prioritized benchmark performance and methodological innovation over in-depth analysis of model failures, biases, and shortcut learning.…
Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation
Volodymyr Havrylov, Haiwen Huang, Dan Zhang +1
Vision Foundation Models (VFMs) are large-scale, pre-trained models that serve as general-purpose backbones for various computer vision tasks. As VFMs' popularity grows, there is a…
StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis
Axel Sauer, Tero Karras, Samuli Laine +2
Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model familie…
Where Does Generative Difficulty Reside? An Empirical Study of Target Representations
Marcel Plocher, Bernhard Schölkopf, Andreas Geiger +1
The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumption is particularly questiona…
Fail2Drive: Benchmarking Closed-Loop Driving Generalization
Simon Gerstenecker, Andreas Geiger, Katrin Renz
Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing b…
GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields
Michael Niemeyer, Andreas Geiger
Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable.…
CaRL: Learning Scalable Planning Policies with Simple Rewards
Bernhard Jaeger, Daniel Dauner, Jens BeiÃwenger +3
We investigate reinforcement learning (RL) for privileged planning in autonomous driving. State-of-the-art approaches for this task are rule-based, but these methods do not scale t…
NeLF-Pro: Neural Light Field Probes for Multi-Scale Novel View Synthesis
Zinuo You, Andreas Geiger, Anpei Chen
We present NeLF-Pro, a novel representation to model and reconstruct light fields in diverse natural scenes that vary in extent and spatial granularity. In contrast to previous fas…
KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs
Christian Reiser, Songyou Peng, Yiyi Liao +1
NeRF synthesizes novel views of a scene with unprecedented quality by fitting a neural radiance field to RGB images. However, NeRF requires querying a deep Multi-Layer Perceptron (…
Unimotion: Unifying 3D Human Motion Synthesis and Understanding
Chuqiao Li, Julian Chibane, Yannan He +3
We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control a…
MERF: Memory-Efficient Radiance Fields for Real-time View Synthesis in Unbounded Scenes
Christian Reiser, Richard Szeliski, Dor Verbin +5
Neural radiance fields enable state-of-the-art photorealistic view synthesis. However, existing radiance field representations are either too compute-intensive for real-time render…
TensoRF: Tensorial Radiance Fields
Anpei Chen, Zexiang Xu, Andreas Geiger +2
We present TensoRF, a novel approach to model and reconstruct radiance fields. Unlike NeRF that purely uses MLPs, we model the radiance field of a scene as a 4D tensor, which repre…
CAMPARI: Camera-Aware Decomposed Generative Neural Radiance Fields
Michael Niemeyer, Andreas Geiger
Tremendous progress in deep generative models has led to photorealistic image synthesis. While achieving compelling results, most approaches operate in the two-dimensional image do…
Efficient Depth-Guided Urban View Synthesis
Sheng Miao, Jiaxin Huang, Dongfeng Bai +4
Recent advances in implicit scene representation enable high-fidelity street view novel view synthesis. However, existing methods optimize a neural radiance field for each scene, r…
VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel Grids
Katja Schwarz, Axel Sauer, Michael Niemeyer +2
State-of-the-art 3D-aware generative models rely on coordinate-based MLPs to parameterize 3D radiance fields. While demonstrating impressive results, querying an MLP for every samp…
NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
Daniel Dauner, Marcel Hallgarten, Tianyu Li +9
Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On th…
GOOD: Exploring Geometric Cues for Detecting Objects in an Open World
Haiwen Huang, Andreas Geiger, Dan Zhang
We address the task of open-world class-agnostic object detection, i.e., detecting every object in an image by learning from a limited number of base object classes. State-of-the-a…
SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes
Xu Chen, Yufeng Zheng, Michael J. Black +2
Neural implicit surface representations have emerged as a promising paradigm to capture 3D shapes in a continuous and resolution-independent manner. However, adapting them to artic…
ATISS: Autoregressive Transformers for Indoor Scene Synthesis
Despoina Paschalidou, Amlan Kar, Maria Shugrina +3
The ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to dat…
LaRa: Efficient Large-Baseline Radiance Fields
Anpei Chen, Haofei Xu, Stefano Esposito +2
Radiance field methods have achieved photorealistic novel view synthesis and geometry reconstruction. But they are mostly applied in per-scene optimization or small-baseline settin…
Category Level Object Pose Estimation via Neural Analysis-by-Synthesis
Xu Chen, Zijian Dong, Jie Song +2
Many object pose estimation algorithms rely on the analysis-by-synthesis framework which requires explicit representations of individual object instances. In this paper we combine…
Learning Unsupervised Hierarchical Part Decomposition of 3D Objects from a Single RGB Image
Despoina Paschalidou, Luc van Gool, Andreas Geiger
Humans perceive the 3D world as a set of distinct objects that are characterized by various low-level (geometry, reflectance) and high-level (connectivity, adjacency, symmetry) pro…
Is Single-View Mesh Reconstruction Ready for Robotics?
Frederik Nolte, Andreas Geiger, Bernhard Schölkopf +1
This paper evaluates single-view mesh reconstruction models for their potential in enabling instant digital twin creation for real-time planning and dynamics prediction using physi…
Renovating Names in Open-Vocabulary Segmentation Benchmarks
Haiwen Huang, Songyou Peng, Dan Zhang +1
Names are essential to both human cognition and vision-language models. Open-vocabulary models utilize class names as text prompts to generalize to categories unseen during trainin…
Learning Cascaded Detection Tasks with Weakly-Supervised Domain Adaptation
Niklas Hanselmann, Nick Schneider, Benedikt Ortelt +1
In order to handle the challenges of autonomous driving, deep learning has proven to be crucial in tackling increasingly complex tasks, such as 3D detection or instance segmentatio…
Relightable Full-Body Gaussian Codec Avatars
Shaofei Wang, Tomas Simon, Igor Santesteban +15
We propose Relightable Full-Body Gaussian Codec Avatars, a new approach for modeling relightable full-body avatars with fine-grained details including face and hands. The unique ch…
Towards Unsupervised Learning of Generative Models for 3D Controllable Image Synthesis
Yiyi Liao, Katja Schwarz, Lars Mescheder +1
In recent years, Generative Adversarial Networks have achieved impressive results in photorealistic image synthesis. This progress nurtures hopes that one day the classical renderi…
2D Gaussian Splatting for Geometrically Accurate Radiance Fields
Binbin Huang, Zehao Yu, Anpei Chen +2
3D Gaussian Splatting (3DGS) has recently revolutionized radiance field reconstruction, achieving high quality novel view synthesis and fast rendering speed without baking. However…
HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
Hongyu Zhou, Jiahao Shao, Lu Xu +6
Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel vie…
123D: Unifying Multi-Modal Autonomous Driving Data at Scale
Daniel Dauner, Valentin Charraut, Bastian Berle +10
The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain largely untapped. Each dataset…
AG3D: Learning to Generate 3D Avatars from 2D Image Collections
Zijian Dong, Xu Chen, Jinlong Yang +3
While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing meth…
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective
Weijie Wang, Qihang Cao, Sensen Gao +10
Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and interacting with the physical…
Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
Xingyu Chen, Yue Chen, Yuliang Xiu +2
Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct super…
Which Training Methods for GANs do actually Converge?
Lars Mescheder, Andreas Geiger, Sebastian Nowozin
Recent work has shown local convergence of GAN training for absolutely continuous data and generator distributions. In this paper, we show that the requirement of absolute continui…
Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes
Zehao Yu, Torsten Sattler, Andreas Geiger
Recently, 3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis results, while allowing the rendering of high-resolution images in real-time. However, lever…
MuRF: Multi-Baseline Radiance Fields
Haofei Xu, Anpei Chen, Yuedong Chen +5
We present Multi-Baseline Radiance Fields (MuRF), a general feed-forward approach to solving sparse view synthesis under multiple different baseline settings (small and large basel…
Mip-Splatting: Alias-free 3D Gaussian Splatting
Zehao Yu, Anpei Chen, Binbin Huang +2
Recently, 3D Gaussian Splatting has demonstrated impressive novel view synthesis results, reaching high fidelity and efficiency. However, strong artifacts can be observed when chan…
SMD-Nets: Stereo Mixture Density Networks
Fabio Tosi, Yiyi Liao, Carolin Schmitt +1
Despite stereo matching accuracy has greatly improved by deep learning in the last few years, recovering sharp boundaries and high-resolution outputs efficiently remains challengin…
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
Peizheng Li, Zhenghao Zhang, David Holtz +6
End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning ca…
LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
Long Nguyen, Micha Fauth, Bernhard Jaeger +4
Simulators can generate virtually unlimited driving data, yet imitation learning policies in simulation still struggle to achieve robust closed-loop performance. Motivated by this…
GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs
Gege Gao, Weiyang Liu, Anpei Chen +2
As pretrained text-to-image diffusion models become increasingly powerful, recent efforts have been made to distill knowledge from these text-to-image pretrained models for optimiz…
PINA: Learning a Personalized Implicit Neural Avatar from a Single RGB-D Video Sequence
Zijian Dong, Chen Guo, Jie Song +3
We present a novel method to learn Personalized Implicit Neural Avatars (PINA) from a short RGB-D sequence. This allows non-expert users to create a detailed and personalized virtu…
Shape As Points: A Differentiable Poisson Solver
Songyou Peng, Chiyu "Max" Jiang, Yiyi Liao +3
In recent years, neural implicit representations gained popularity in 3D reconstruction due to their expressiveness and flexibility. However, the implicit nature of neural implicit…
Towards Scalable Multi-View Reconstruction of Geometry and Materials
Carolin Schmitt, Božidar AntiÄ, Andrei Neculai +2
In this paper, we propose a novel method for joint recovery of camera pose, object geometry and spatially-varying Bidirectional Reflectance Distribution Function (svBRDF) of 3D sce…
Centaur: Robust End-to-End Autonomous Driving with Test-Time Training
Chonghao Sima, Kashyap Chitta, Zhiding Yu +5
How can we rely on an end-to-end autonomous vehicle's complex decision-making system during deployment? One common solution is to have a ``fallback layer'' that checks the planned…
Learning 3D Shape Completion under Weak Supervision
David Stutz, Andreas Geiger
We address the problem of 3D shape completion from sparse and noisy point clouds, a fundamental problem in computer vision and robotics. Recent approaches are either data-driven or…
Augmented Reality Meets Computer Vision : Efficient Data Generation for Urban Driving Scenes
Hassan Abu Alhaija, Siva Karthik Mustikovela, Lars Mescheder +2
The success of deep learning in computer vision is based on availability of large annotated datasets. To lower the need for hand labeled images, virtually rendered 3D worlds have r…
EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis
Sheng Miao, Sijin Li, Pan Wang +5
Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time wit…
An Invitation to Deep Reinforcement Learning
Bernhard Jaeger, Andreas Geiger
Training a deep neural network to maximize a target objective has become the standard recipe for successful machine learning over the last decade. These networks can be optimized w…
HDT: Hierarchical Document Transformer
Haoyu He, Markus Flicke, Jan Buchmann +2
In this paper, we propose the Hierarchical Document Transformer (HDT), a novel sparse Transformer architecture tailored for structured hierarchical documents. Such documents are ex…
Convolutional Occupancy Networks
Songyou Peng, Michael Niemeyer, Lars Mescheder +2
Recently, implicit neural representations have gained popularity for learning-based 3D reconstruction. While demonstrating promising results, most implicit approaches are limited t…
Learning Implicit Surface Light Fields
Michael Oechsle, Michael Niemeyer, Lars Mescheder +2
Implicit representations of 3D objects have recently achieved impressive results on learning-based 3D reconstruction tasks. While existing works use simple texture models to repres…
InvSplat: Inverse Feed-Forward Scene Splatting
Polina Karpikova, Wenjing Bian, Haofei Xu +2
Inverse rendering aims to recover both 3D geometry and physically meaningful material properties from images, enabling applications such as relighting and novel view synthesis. Opt…
Benchmarking Unsupervised Object Representations for Video Sequences
Marissa A. Weis, Kashyap Chitta, Yash Sharma +4
Perceiving the world in terms of objects and tracking them through time is a crucial prerequisite for reasoning and scene understanding. Recently, several methods have been propose…
Intrinsic Autoencoders for Joint Neural Rendering and Intrinsic Image Decomposition
Hassan Abu Alhaija, Siva Karthik Mustikovela, Justus Thies +4
Neural rendering techniques promise efficient photo-realistic image synthesis while at the same time providing rich control over scene parameters by learning the physical image for…
DepthSplat: Connecting Gaussian Splatting and Depth
Haofei Xu, Songyou Peng, Fangjinhua Wang +4
Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and…
Slots, Transitions, Loops: Learning Composable World Models for ARC
Gege Gao, Bernhard Schölkopf, Andreas Geiger
ARC tests in-context rule induction: given a few input-output demonstrations, a model must infer the hidden rule and apply it to a new query. While many approaches express ARC rule…
GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis
Katja Schwarz, Yiyi Liao, Michael Niemeyer +1
While 2D generative adversarial networks have enabled high-resolution image synthesis, they largely lack an understanding of the 3D world and the image formation process. Thus, the…
Self-Supervised Linear Motion Deblurring
Peidong Liu, Joel Janai, Marc Pollefeys +2
Motion blurry images challenge many computer vision algorithms, e.g, feature detection, motion estimation, or object recognition. Deep convolutional neural networks are state-of-th…
Volumetric Surfaces: Representing Fuzzy Geometries with Layered Meshes
Stefano Esposito, Anpei Chen, Christian Reiser +7
High-quality view synthesis relies on volume rendering, splatting, or surface rendering. While surface rendering is typically the fastest, it struggles to accurately model fuzzy ge…
Artificial Kuramoto Oscillatory Neurons
Takeru Miyato, Sindy Löwe, Andreas Geiger +1
It has long been known in both neuroscience and AI that ``binding'' between neurons leads to a form of competitive learning where representations are compressed in order to represe…
RayNet: Learning Volumetric 3D Reconstruction with Ray Potentials
Despoina Paschalidou, Ali Osman Ulusoy, Carolin Schmitt +2
In this paper, we consider the problem of reconstructing a dense 3D model using images captured from different views. Recent methods based on convolutional neural networks (CNN) al…
Learn2Splat: Extending the Horizon of Learned 3DGS Optimization
Naama Pearl, Stefano Esposito, Haofei Xu +6
3D Gaussian Splatting (3DGS) optimization is most commonly performed using standard optimizers (Adam, SGD). While stable across diverse scenes, standard optimizers are general-purp…
sshELF: Single-Shot Hierarchical Extrapolation of Latent Features for 3D Reconstruction from Sparse-Views
Eyvaz Najafli, Marius Kästingschäfer, Sebastian Bernhard +2
Reconstructing unbounded outdoor scenes from sparse outward-facing views poses significant challenges due to minimal view overlap. Previous methods often lack cross-scene understan…
KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D
Yiyi Liao, Jun Xie, Andreas Geiger
For the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other…
MOTS: Multi-Object Tracking and Segmentation
Paul Voigtlaender, Michael Krause, Aljosa Osep +4
This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two ex…
Unifying Flow, Stereo and Depth Estimation
Haofei Xu, Jing Zhang, Jianfei Cai +4
We present a unified formulation and model for three motion and 3D perception tasks: optical flow, rectified stereo matching and unrectified stereo depth estimation from posed imag…
Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
Merve Kocabas, Gege Gao, Bernhard Schölkopf +1
Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, wh…
RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs
Michael Niemeyer, Jonathan T. Barron, Ben Mildenhall +3
Neural Radiance Fields (NeRF) have emerged as a powerful representation for the task of novel view synthesis due to their simplicity and state-of-the-art performance. Though NeRF c…
PlanT: Explainable Planning Transformers via Object-Level Representations
Katrin Renz, Kashyap Chitta, Otniel-Bogdan Mercea +3
Planning an optimal route in a complex environment requires efficient reasoning about the surrounding scene. While human drivers prioritize important objects and ignore details not…
HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving
Hongyu Zhou, Longzhong Lin, Jiabao Wang +6
In the past few decades, autonomous driving algorithms have made significant progress in perception, planning, and control. However, evaluating individual components does not fully…
MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction
Zijian Dong, Longteng Duan, Jie Song +2
We present MoGA, a novel method to reconstruct high-fidelity 3D Gaussian avatars from a single-view image. The main challenge lies in inferring unseen appearance and geometric deta…
Learning Neural Light Transport
Paul Sanzenbacher, Lars Mescheder, Andreas Geiger
In recent years, deep generative models have gained significance due to their ability to synthesize natural-looking images with applications ranging from virtual reality to data au…
Hidden Biases of End-to-End Driving Datasets
Julian Zimmerlin, Jens BeiÃwenger, Bernhard Jaeger +2
End-to-end driving systems have made rapid progress, but have so far not been applied to the challenging new CARLA Leaderboard 2.0. Further, while there is a large body of literatu…
Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle +1
Learning-based 3D reconstruction methods have shown impressive results. However, most methods require 3D supervision which is often hard to obtain for real-world datasets. Recently…
World Engine: Towards the Era of Post-Training for Autonomous Driving
Tianyu Li, Li Chen, Caojun Wang +16
Autonomous vehicles must operate safely in the real world, where errors can have severe consequences. Although modern end-to-end driving policies excel in routine scenarios, their…
On the Integration of Optical Flow and Action Recognition
Laura Sevilla-Lara, Yiyi Liao, Fatma Guney +3
Most of the top performing action recognition methods use optical flow as a "black box" input. Here we take a deeper look at the combination of flow and action recognition, and inv…
3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting
Zhiyin Qian, Shaofei Wang, Marko Mihajlovic +2
We introduce an approach that creates animatable human avatars from monocular videos using 3D Gaussian Splatting (3DGS). Existing methods based on neural radiance fields (NeRFs) ac…
Label Efficient Visual Abstractions for Autonomous Driving
Aseem Behl, Kashyap Chitta, Aditya Prakash +2
It is well known that semantic segmentation can be used as an effective intermediate representation for learning driving policies. However, the task of street scene semantic segmen…
Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
Aditya Prakash, Kashyap Chitta, Andreas Geiger
How should representations from complementary sensors be integrated for autonomous driving? Geometry-based sensor fusion has shown great promise for perception tasks such as object…
gDNA: Towards Generative Detailed Neural Avatars
Xu Chen, Tianjian Jiang, Jie Song +4
To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging…
Counterfactual Generative Networks
Axel Sauer, Andreas Geiger
Neural networks are prone to learning shortcuts -- they often model simple correlations, ignoring more complex ones that potentially generalize better. Prior works on image classif…