papers

Publications (46)

cs.CV2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

Junxuan Li, Rawal Khirodkar, Chengan He +37

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans wit…

cs.GR2022

Dressing Avatars: Deep Photorealistic Appearance for Physically Simulated Clothing

Donglai Xiang, Timur Bagautdinov, Tuur Stuyck +10

Despite recent progress in developing animatable full-body avatars, realistic modeling of clothing - one of the core aspects of human self-expression - remains an open challenge. S…

cs.CV2019

Shapes and Context: In-the-Wild Image Synthesis & Manipulation

Aayush Bansal, Yaser Sheikh, Deva Ramanan

We introduce a data-driven approach for interactively synthesizing in-the-wild images from semantic label maps. Our approach is dramatically different from recent work in this spac…

cs.CV2018

Structure from Recurrent Motion: From Rigidity to Recurrency

Xiu Li, Hongdong Li, Hanbyul Joo +2

This paper proposes a new method for Non-Rigid Structure-from-Motion (NRSfM) from a long monocular video sequence observing a non-rigid object performing recurrent and possibly rep…

cs.GR2018

Deep Appearance Models for Face Rendering

Stephen Lombardi, Jason Saragih, Tomas Simon +1

We introduce a deep appearance model for rendering the human face. Inspired by Active Appearance Models, we develop a data-driven rendering pipeline that learns a joint representat…

cs.CY2018

Video Analysis for Body-worn Cameras in Law Enforcement

Jason J. Corso, Alexandre Alahi, Kristen Grauman +4

The social conventions and expectations around the appropriate use of imaging and video has been transformed by the availability of video cameras in our pockets. The impact on law…

cs.CV2021

High-fidelity Face Tracking for AR/VR via Deep Lighting Adaptation

Lele Chen, Chen Cao, Fernando De la Torre +3

3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven…

cs.CV2019

To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations

Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency +1

Non verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepr…

cs.CV2016

How useful is photo-realistic rendering for visual learning?

Yair Movshovitz-Attias, Takeo Kanade, Yaser Sheikh

Data seems cheap to get, and in many ways it is, but the process of creating a high quality labeled dataset from a mass of data is time-consuming and expensive. With the advent of…

cs.CV2019

LBS Autoencoder: Self-supervised Fitting of Articulated Meshes to Point Clouds

Chun-Liang Li, Tomas Simon, Jason Saragih +2

We present LBS-AE; a self-supervised autoencoding algorithm for fitting articulated mesh models to point clouds. As input, we take a sequence of point clouds to be registered as we…

cs.CV2016

Convolutional Pose Machines

Shih-En Wei, Varun Ramakrishna, Takeo Kanade +1

Pose Machines provide a sequential prediction framework for learning rich implicit spatial models. In this work we show a systematic design for how convolutional networks can be in…

cs.CV2018

Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies

Hanbyul Joo, Tomas Simon, Yaser Sheikh

We present a unified deformation model for the markerless capture of multiple scales of human movement, including facial expressions, body motion, and hand gestures. An initial mod…

cs.CV2023

Multiface: A Dataset for Neural Face Rendering

Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson +28

Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering…

cs.CV2019

Single-Network Whole-Body Pose Estimation

Gines Hidalgo, Yaadhav Raaj, Haroon Idrees +4

We present the first single-network approach for 2D~whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-…

cs.CV2017

Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

Zhe Cao, Tomas Simon, Shih-En Wei +1

We present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields…

cs.CV2016

Panoptic Studio: A Massively Multiview System for Social Interaction Capture

Hanbyul Joo, Tomas Simon, Xulong Li +10

We present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is fun…

cs.CV2021

Pixel Codec Avatars

Shugao Ma, Tomas Simon, Jason Saragih +4

Telecommunication with photorealistic avatars in virtual or augmented reality is a promising path for achieving authentic face-to-face communication in 3D over remote physical dist…

cs.CV2018

Recycle-GAN: Unsupervised Video Retargeting

Aayush Bansal, Shugao Ma, Deva Ramanan +1

We introduce a data-driven approach for unsupervised video retargeting that translates content from one domain to another while preserving the style native to a domain, i.e., if co…

cs.CV2025

FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images

Rong Wang, Fabian Prada, Ziyan Wang +10

We present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and c…

cs.CV2019

Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in A Triadic Interaction

Hanbyul Joo, Tomas Simon, Mina Cikara +1

We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a…

cs.CV2021

Driving-Signal Aware Full-Body Avatars

Timur Bagautdinov, Chenglei Wu, Tomas Simon +6

We present a learning-based method for building driving-signal aware full-body avatars. Our model is a conditional variational autoencoder that can be animated with incomplete driv…

cs.CV2020

Spatiotemporal Bundle Adjustment for Dynamic 3D Human Reconstruction in the Wild

Minh Vo, Yaser Sheikh, Srinivasa G. Narasimhan

Bundle adjustment jointly optimizes camera intrinsics and extrinsics and 3D point triangulation to reconstruct a static scene. The triangulation constraint, however, is invalid for…

cs.GR2019

Neural Volumes: Learning Dynamic Renderable Volumes from Images

Stephen Lombardi, Tomas Simon, Jason Saragih +3

Modeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occ…

cs.CV2020

Self-supervised Multi-view Person Association and Its Applications

Minh Vo, Ersin Yumer, Kalyan Sunkavalli +3

Reliable markerless motion tracking of people participating in a complex group activity from multiple moving cameras is challenging due to frequent occlusions, strong viewpoint and…

cs.CV2024

Universal Facial Encoding of Codec Avatars from VR Headsets

Shaojie Bai, Te-Li Wang, Chenghui Li +8

Faithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficien…

cs.CV2018

Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors

Xuanyi Dong, Shoou-I Yu, Xinshuo Weng +3

In this paper, we present supervision-by-registration, an unsupervised approach to improve the precision of facial landmark detectors on both images and video. Our key observation…

cs.CV2022

MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement

Alexander Richard, Michael Zollhoefer, Yandong Wen +2

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face…

cs.CV2019

Efficient Online Multi-Person 2D Pose Tracking with Recurrent Spatio-Temporal Affinity Fields

Yaadhav Raaj, Haroon Idrees, Gines Hidalgo +1

We present an online approach to efficiently and simultaneously detect and track the 2D pose of multiple people in a video sequence. We build upon Part Affinity Field (PAF) represe…

cs.CV2018

Monocular Total Capture: Posing Face, Body, and Hands in the Wild

Donglai Xiang, Hanbyul Joo, Yaser Sheikh

We present the first method to capture the 3D total motion of a target person from a monocular view input. Given an image or a monocular video, our method reconstructs the motion f…

cs.CV2021

Supervision by Registration and Triangulation for Landmark Detection

Xuanyi Dong, Yi Yang, Shih-En Wei +3

We present Supervision by Registration and Triangulation (SRT), an unsupervised approach that utilizes unlabeled multi-view video to improve the accuracy and precision of landmark…

cs.CV2017

Hand Keypoint Detection in Single Images using Multiview Bootstrapping

Tomas Simon, Hanbyul Joo, Iain Matthews +1

We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedu…

cs.CV2024

URAvatar: Universal Relightable Gaussian Codec Avatars

Junxuan Li, Chen Cao, Gabriel Schwartz +5

We present a new approach to creating photorealistic and relightable head avatars from a phone scan with unknown illumination. The reconstructed avatars can be animated and relit i…

cs.CV2023

RelightableHands: Efficient Neural Relighting of Articulated Hand Models

Shun Iwase, Shunsuke Saito, Tomas Simon +7

We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a tea…

cs.CV2020

Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels

Yi Zhou, Chenglei Wu, Zimo Li +5

Learning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher pre…

cs.GR2021

Mixture of Volumetric Primitives for Efficient Neural Rendering

Stephen Lombardi, Tomas Simon, Gabriel Schwartz +3

Real-time rendering and animation of humans is a core function in games, movies, and telepresence applications. Existing methods have a number of drawbacks we aim to address with o…

cs.CV2025

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

Chen Guo, Junxuan Li, Yash Kant +3

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animati…

cs.CV2020

Expressive Telepresence via Modular Codec Avatars

Hang Chu, Shugao Ma, Fernando De la Torre +2

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow vide…

cs.CV2022

Garment Avatars: Realistic Cloth Driving using Pattern Registration

Oshri Halimi, Fabian Prada, Tuur Stuyck +7

Virtual telepresence is the future of online communication. Clothing is an essential part of a person's identity and self-expression. Yet, ground truth data of registered clothes i…

cs.CV2022

Drivable Volumetric Avatars using Texel-Aligned Features

Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito +8

Photorealistic telepresence requires both high-fidelity body modeling and faithful driving to enable dynamically synthesized appearance that is indistinguishable from reality. In t…

cs.CV2018

Capture Dense: Markerless Motion Capture Meets Dense Pose Estimation

Xiu Li, Yebin Liu, Hanbyul Joo +2

We present a method to combine markerless motion capture and dense pose feature estimation into a single framework. We demonstrate that dense pose information can help for multivie…

cs.CV2017

PixelNN: Example-based Image Synthesis

Aayush Bansal, Yaser Sheikh, Deva Ramanan

We present a simple nearest-neighbor (NN) approach that synthesizes high-frequency photorealistic images from an "incomplete" signal such as a low-resolution image, a surface norma…

cs.CV2024

Rasterized Edge Gradients: Handling Discontinuities Differentiably

Stanislav Pidhorskyi, Tomas Simon, Gabriel Schwartz +3

Computing the gradients of a rendering process is paramount for diverse applications in computer vision and graphics. However, accurate computation of these gradients is challengin…

cs.CV2020

Audio- and Gaze-driven Facial Animation of Codec Avatars

Alexander Richard, Colin Lea, Shugao Ma +3

Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are al…

cs.CV2019

OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

Zhe Cao, Gines Hidalgo, Tomas Simon +2

Realtime multi-person 2D pose estimation is a key component in enabling machines to have an understanding of people in images and videos. In this work, we present a realtime approa…

cs.CV2024

URHand: Universal Relightable Hands

Zhaoxi Chen, Gyeongsik Moon, Kaiwen Guo +20

Existing photorealistic relightable hand models require extensive identity-specific observations in different views, poses, and illuminations, and face challenges in generalizing t…

cs.CV2020

4D Visualization of Dynamic Events from Unconstrained Multi-View Videos

Aayush Bansal, Minh Vo, Yaser Sheikh +2

We present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-super…