collaborators

9 papers

cs.CV2026

SkillSight: Efficient First-Person Skill Assessment with Gaze

Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman

Egocentric perception on smart glasses could transform how we learn new skills in the physical world, but automatic skill assessment remains a fundamental technical challenge. We i…

cs.CV2026

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions

Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman

When obtaining visual illustrations from text descriptions, today's methods take a description with a single text context - a caption, or an action description - and retrieve or ge…

cs.CV2026

SportSkills: Physical Skill Learning from Sports Instructional Videos

Kumar Ashutosh, Chi Hsuan Wu, Kristen Grauman

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce Sp…

cs.CV2026

Human detectors are surprisingly powerful reward models

Kumar Ashutosh, XuDong Wang, Xi Yin +4

Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…

cs.CV2025

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual spor…

cs.HC2025

Vid2Coach: Transforming How-To Videos into Task Assistants

Mina Huh, Zihui Xue, Ujjaini Das +3

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our o…