most citedSocial Scene Understanding: End-to-End Multi-Person Action Localization and Collective Activity Recognition

7 citations · 7 across the 1 of their papers we have counts for

collaborators

6 papers

cs.CV20241 cited

Sapiens: Foundation for Human Vision Models

Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez +5

We present Sapiens, a family of models for four fundamental human-centric vision tasks -- 2D pose estimation, body-part segmentation, depth estimation, and surface normal predictio…

cs.CV2024

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

Evonne Ng, Javier Romero, Timur Bagautdinov +4

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we outpu…

cs.GR202316 cited

Drivable Avatar Clothing: Faithful Full-Body Telepresence with Dynamic Clothing Driven by Sparse RGB-D Input

Donglai Xiang, Fabian Prada, Zhe Cao +4

Clothing is an important part of human appearance but challenging to model in photorealistic avatars. In this work we present avatars with dynamically moving loose clothing that ca…

cs.CV20231 cited

RelightableHands: Efficient Neural Relighting of Articulated Hand Models

Shun Iwase, Shunsuke Saito, Tomas Simon +7

We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a tea…

cs.CV20221 cited

Drivable Volumetric Avatars using Texel-Aligned Features

Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito +8

Photorealistic telepresence requires both high-fidelity body modeling and faithful driving to enable dynamically synthesized appearance that is indistinguishable from reality. In t…

cs.CV20167 cited

Social Scene Understanding: End-to-End Multi-Person Action Localization and Collective Activity Recognition

Timur Bagautdinov, Alexandre Alahi, François Fleuret +2

We present a unified framework for understanding human social behaviors in raw image sequences. Our model jointly detects multiple individuals, infers their social actions, and est…