9 papers
ReGenHuman: Re-Generating Human Appearances for Realistic Full-Body Video Anonymization
Adam Sun, Eshaan Barkataki, Arnold Milstein +2
Anonymizing human-centric video data is an understudied problem. Prior anonymization techniques either blur or redact pixels at the cost of realism and downstream utility, or gener…
ESAM++: Efficient Online 3D Perception on the Edge
Qin Liu, Lavisha Aggarwal, Saptarashmi Bandyopadhyay +4
Online 3D scene perception in real time is essential for robotics, AR/VR, and autonomous systems, particularly in edge computing scenarios where computational resources are limited…
AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion
Hongjie Li, Heng Yu, Jiaman Li +4
Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of human behavior. Existing metho…
GenFusion: Feed-forward Human Performance Capture via Progressive Canonical Space Updates
Youngjoong Kwon, Yao He, Heejung Choi +4
We present a feed-forward human performance capture method that renders novel views of a performer from a monocular RGB stream. A key challenge in this setting is the lack of suffi…
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi +3
Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologie…
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Noriyuki Kugo, Xiang Li, Zixin Li +9
Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. Howe…