9 papers
SkillSight: Efficient First-Person Skill Assessment with Gaze
Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman
Egocentric perception on smart glasses could transform how we learn new skills in the physical world, but automatic skill assessment remains a fundamental technical challenge. We i…
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman
When obtaining visual illustrations from text descriptions, today's methods take a description with a single text context - a caption, or an action description - and retrieve or ge…
SportSkills: Physical Skill Learning from Sports Instructional Videos
Kumar Ashutosh, Chi Hsuan Wu, Kristen Grauman
Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce Sp…
Human detectors are surprisingly powerful reward models
Kumar Ashutosh, XuDong Wang, Xi Yin +4
Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…
Learning Skill-Attributes for Transferable Assessment in Video
Kumar Ashutosh, Kristen Grauman
Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual spor…
Vid2Coach: Transforming How-To Videos into Task Assistants
Mina Huh, Zihui Xue, Ujjaini Das +3
People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our o…