8 papers · 1 filter
SkillSight: Efficient First-Person Skill Assessment with Gaze
Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman
Egocentric perception on smart glasses could transform how we learn new skills in the physical world, but automatic skill assessment remains a fundamental technical challenge. We i…
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman
When obtaining visual illustrations from text descriptions, today's methods take a description with a single text context - a caption, or an action description - and retrieve or ge…
SportSkills: Physical Skill Learning from Sports Instructional Videos
Kumar Ashutosh, Chi Hsuan Wu, Kristen Grauman
Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce Sp…
Human detectors are surprisingly powerful reward models
Kumar Ashutosh, XuDong Wang, Xi Yin +4
Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…
Learning Skill-Attributes for Transferable Assessment in Video
Kumar Ashutosh, Kristen Grauman
Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual spor…
FIction: 4D Future Interaction Prediction from Video
Kumar Ashutosh, Georgios Pavlakos, Kristen Grauman
Anticipating how a person will interact with objects in an environment is essential for activity understanding, but existing methods are limited to the 2D space of video frames-cap…