papers

Publications (6)

cs.CV2025

Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero +81

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent…

cs.CV2018

Watch to Edit: Video Retargeting using Gaze

Kranthi Kumar, Moneish Kumar, Vineet Gandhi +1

We present a novel approach to optimally retarget videos for varied displays with differing aspect ratios by preserving salient scene content discovered via eye tracking. Our algor…

cs.CV2020

GAZED- Gaze-guided Cinematic Editing of Wide-Angle Monocular Video Recordings

K L Bhanu Moorthy, Moneish Kumar, Ramanathan Subramaniam +1

We present GAZED- eye GAZe-guided EDiting for videos captured by a solitary, static, wide-angle and high-resolution camera. Eye-gaze has been effectively employed in computational…

cs.CV2023

S2RF: Semantically Stylized Radiance Fields

Dishani Lahiri, Neeraj Panse, Moneish Kumar

We present our method for transferring style from any arbitrary image(s) to object(s) within a 3D scene. Our primary objective is to offer more control in 3D scene stylization, fac…

cs.CV2025

Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset

Claire McLean, Makenzie Meendering, Tristan Swartz +21

The Codec Avatars Lab at Meta introduces Embody 3D, a multimodal dataset of 500 individual hours of 3D motion data from 439 participants collected in a multi-camera collection stag…

cs.CV2024

Cameras as Rays: Pose Estimation via Ray Diffusion

Jason Y. Zhang, Amy Lin, Moneish Kumar +3

Estimating camera poses is a fundamental task for 3D reconstruction and remains challenging given sparsely sampled views (<10). In contrast to existing approaches that pursue top-d…