Publications (46)
Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
Junxuan Li, Rawal Khirodkar, Chengan He +37
High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans wit…
Dressing Avatars: Deep Photorealistic Appearance for Physically Simulated Clothing
Donglai Xiang, Timur Bagautdinov, Tuur Stuyck +10
Despite recent progress in developing animatable full-body avatars, realistic modeling of clothing - one of the core aspects of human self-expression - remains an open challenge. S…
Shapes and Context: In-the-Wild Image Synthesis & Manipulation
Aayush Bansal, Yaser Sheikh, Deva Ramanan
We introduce a data-driven approach for interactively synthesizing in-the-wild images from semantic label maps. Our approach is dramatically different from recent work in this spac…
Structure from Recurrent Motion: From Rigidity to Recurrency
Xiu Li, Hongdong Li, Hanbyul Joo +2
This paper proposes a new method for Non-Rigid Structure-from-Motion (NRSfM) from a long monocular video sequence observing a non-rigid object performing recurrent and possibly rep…
Deep Appearance Models for Face Rendering
Stephen Lombardi, Jason Saragih, Tomas Simon +1
We introduce a deep appearance model for rendering the human face. Inspired by Active Appearance Models, we develop a data-driven rendering pipeline that learns a joint representat…
Video Analysis for Body-worn Cameras in Law Enforcement
Jason J. Corso, Alexandre Alahi, Kristen Grauman +4
The social conventions and expectations around the appropriate use of imaging and video has been transformed by the availability of video cameras in our pockets. The impact on law…
High-fidelity Face Tracking for AR/VR via Deep Lighting Adaptation
Lele Chen, Chen Cao, Fernando De la Torre +3
3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven…
To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations
Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency +1
Non verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepr…
How useful is photo-realistic rendering for visual learning?
Yair Movshovitz-Attias, Takeo Kanade, Yaser Sheikh
Data seems cheap to get, and in many ways it is, but the process of creating a high quality labeled dataset from a mass of data is time-consuming and expensive. With the advent of…
LBS Autoencoder: Self-supervised Fitting of Articulated Meshes to Point Clouds
Chun-Liang Li, Tomas Simon, Jason Saragih +2
We present LBS-AE; a self-supervised autoencoding algorithm for fitting articulated mesh models to point clouds. As input, we take a sequence of point clouds to be registered as we…
Convolutional Pose Machines
Shih-En Wei, Varun Ramakrishna, Takeo Kanade +1
Pose Machines provide a sequential prediction framework for learning rich implicit spatial models. In this work we show a systematic design for how convolutional networks can be in…
Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies
Hanbyul Joo, Tomas Simon, Yaser Sheikh
We present a unified deformation model for the markerless capture of multiple scales of human movement, including facial expressions, body motion, and hand gestures. An initial mod…
Multiface: A Dataset for Neural Face Rendering
Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson +28
Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering…
Single-Network Whole-Body Pose Estimation
Gines Hidalgo, Yaadhav Raaj, Haroon Idrees +4
We present the first single-network approach for 2D~whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-…
Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
Zhe Cao, Tomas Simon, Shih-En Wei +1
We present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields…
Panoptic Studio: A Massively Multiview System for Social Interaction Capture
Hanbyul Joo, Tomas Simon, Xulong Li +10
We present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is fun…
Pixel Codec Avatars
Shugao Ma, Tomas Simon, Jason Saragih +4
Telecommunication with photorealistic avatars in virtual or augmented reality is a promising path for achieving authentic face-to-face communication in 3D over remote physical dist…
Recycle-GAN: Unsupervised Video Retargeting
Aayush Bansal, Shugao Ma, Deva Ramanan +1
We introduce a data-driven approach for unsupervised video retargeting that translates content from one domain to another while preserving the style native to a domain, i.e., if co…
FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images
Rong Wang, Fabian Prada, Ziyan Wang +10
We present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and c…
Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in A Triadic Interaction
Hanbyul Joo, Tomas Simon, Mina Cikara +1
We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a…
Driving-Signal Aware Full-Body Avatars
Timur Bagautdinov, Chenglei Wu, Tomas Simon +6
We present a learning-based method for building driving-signal aware full-body avatars. Our model is a conditional variational autoencoder that can be animated with incomplete driv…
Spatiotemporal Bundle Adjustment for Dynamic 3D Human Reconstruction in the Wild
Minh Vo, Yaser Sheikh, Srinivasa G. Narasimhan
Bundle adjustment jointly optimizes camera intrinsics and extrinsics and 3D point triangulation to reconstruct a static scene. The triangulation constraint, however, is invalid for…
Neural Volumes: Learning Dynamic Renderable Volumes from Images
Stephen Lombardi, Tomas Simon, Jason Saragih +3
Modeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occ…
Self-supervised Multi-view Person Association and Its Applications
Minh Vo, Ersin Yumer, Kalyan Sunkavalli +3
Reliable markerless motion tracking of people participating in a complex group activity from multiple moving cameras is challenging due to frequent occlusions, strong viewpoint and…
Universal Facial Encoding of Codec Avatars from VR Headsets
Shaojie Bai, Te-Li Wang, Chenghui Li +8
Faithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficien…
Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
Xuanyi Dong, Shoou-I Yu, Xinshuo Weng +3
In this paper, we present supervision-by-registration, an unsupervised approach to improve the precision of facial landmark detectors on both images and video. Our key observation…
MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
Alexander Richard, Michael Zollhoefer, Yandong Wen +2
This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face…
Efficient Online Multi-Person 2D Pose Tracking with Recurrent Spatio-Temporal Affinity Fields
Yaadhav Raaj, Haroon Idrees, Gines Hidalgo +1
We present an online approach to efficiently and simultaneously detect and track the 2D pose of multiple people in a video sequence. We build upon Part Affinity Field (PAF) represe…
Monocular Total Capture: Posing Face, Body, and Hands in the Wild
Donglai Xiang, Hanbyul Joo, Yaser Sheikh
We present the first method to capture the 3D total motion of a target person from a monocular view input. Given an image or a monocular video, our method reconstructs the motion f…
Supervision by Registration and Triangulation for Landmark Detection
Xuanyi Dong, Yi Yang, Shih-En Wei +3
We present Supervision by Registration and Triangulation (SRT), an unsupervised approach that utilizes unlabeled multi-view video to improve the accuracy and precision of landmark…
Hand Keypoint Detection in Single Images using Multiview Bootstrapping
Tomas Simon, Hanbyul Joo, Iain Matthews +1
We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedu…
URAvatar: Universal Relightable Gaussian Codec Avatars
Junxuan Li, Chen Cao, Gabriel Schwartz +5
We present a new approach to creating photorealistic and relightable head avatars from a phone scan with unknown illumination. The reconstructed avatars can be animated and relit i…
RelightableHands: Efficient Neural Relighting of Articulated Hand Models
Shun Iwase, Shunsuke Saito, Tomas Simon +7
We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a tea…
Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels
Yi Zhou, Chenglei Wu, Zimo Li +5
Learning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher pre…
Mixture of Volumetric Primitives for Efficient Neural Rendering
Stephen Lombardi, Tomas Simon, Gabriel Schwartz +3
Real-time rendering and animation of humans is a core function in games, movies, and telepresence applications. Existing methods have a number of drawbacks we aim to address with o…
Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior
Chen Guo, Junxuan Li, Yash Kant +3
We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animati…
Expressive Telepresence via Modular Codec Avatars
Hang Chu, Shugao Ma, Fernando De la Torre +2
VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow vide…
Garment Avatars: Realistic Cloth Driving using Pattern Registration
Oshri Halimi, Fabian Prada, Tuur Stuyck +7
Virtual telepresence is the future of online communication. Clothing is an essential part of a person's identity and self-expression. Yet, ground truth data of registered clothes i…
Drivable Volumetric Avatars using Texel-Aligned Features
Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito +8
Photorealistic telepresence requires both high-fidelity body modeling and faithful driving to enable dynamically synthesized appearance that is indistinguishable from reality. In t…
Capture Dense: Markerless Motion Capture Meets Dense Pose Estimation
Xiu Li, Yebin Liu, Hanbyul Joo +2
We present a method to combine markerless motion capture and dense pose feature estimation into a single framework. We demonstrate that dense pose information can help for multivie…
PixelNN: Example-based Image Synthesis
Aayush Bansal, Yaser Sheikh, Deva Ramanan
We present a simple nearest-neighbor (NN) approach that synthesizes high-frequency photorealistic images from an "incomplete" signal such as a low-resolution image, a surface norma…
Rasterized Edge Gradients: Handling Discontinuities Differentiably
Stanislav Pidhorskyi, Tomas Simon, Gabriel Schwartz +3
Computing the gradients of a rendering process is paramount for diverse applications in computer vision and graphics. However, accurate computation of these gradients is challengin…
Audio- and Gaze-driven Facial Animation of Codec Avatars
Alexander Richard, Colin Lea, Shugao Ma +3
Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are al…
OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
Zhe Cao, Gines Hidalgo, Tomas Simon +2
Realtime multi-person 2D pose estimation is a key component in enabling machines to have an understanding of people in images and videos. In this work, we present a realtime approa…
URHand: Universal Relightable Hands
Zhaoxi Chen, Gyeongsik Moon, Kaiwen Guo +20
Existing photorealistic relightable hand models require extensive identity-specific observations in different views, poses, and illuminations, and face challenges in generalizing t…
4D Visualization of Dynamic Events from Unconstrained Multi-View Videos
Aayush Bansal, Minh Vo, Yaser Sheikh +2
We present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-super…