collaborators

8 papers

cs.CV2025

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara +1

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but…

cs.CV2025

Mask2IV: Interaction-Centric Video Generation via Mask Trajectories

Gen Li, Bo Zhao, Jianfei Yang +1

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse vis…

cs.CV2025

Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training

Shriram M Sathiyanarayanan, Xinyue Hao, Shihao Hou +4

The success of the machine learning field has reliably depended on training on large datasets. While effective, this trend comes at an extraordinary cost. This is due to two deeply…

cs.RO2025

Learning Precise Affordances from Egocentric Videos for Robotic Manipulation

Gen Li, Nikolaos Tsagkas, Jifei Song +4

Affordance, defined as the potential actions that an object offers, is crucial for embodied AI agents. For example, such knowledge directs an agent to grasp a knife by the handle f…

cs.CL2025

Predicting Implicit Arguments in Procedural Video Instructions

Anil Batra, Laura Sevilla-Lara, Marcus Rohrbach +1

Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by id…

cs.CV2025

Principles of Visual Tokens for Efficient Video Understanding

Xinyue Hao, Gen Li, Shreyank N Gowda +4

Video understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redu…