papers

Publications (21)

cs.CV2023

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani +1

cs.CV2025

LLMs can see and hear without any training

Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen +2

cs.CV2023

Video-Mined Task Graphs for Keystep Recognition in Instructional Videos

Kumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras +1

cs.CV2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

cs.CV2025

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

cs.LG2020

Bandit algorithms: Letting go of logarithmic regret for statistical robustness

Kumar Ashutosh, Jayakrishnan Nair, Anmol Kagrecha +1

cs.CV2024

Detours for Navigating Instructional Videos

Kumar Ashutosh, Zihui Xue, Tushar Nagarajan +1

cs.CV2022

RoS-KD: A Robust Stochastic Knowledge Distillation Approach for Noisy Medical Imaging

Ajay Jaiswal, Kumar Ashutosh, Justin F Rousseau +3

cs.CV2023

HierVL: Learning Hierarchical Video-Language Embeddings

Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani +1

cs.CV2026

SportSkills: Physical Skill Learning from Sports Instructional Videos

Kumar Ashutosh, Chi Hsuan Wu, Kristen Grauman

cs.LG2020

Lower Bounds for Policy Iteration on Multi-action MDPs

Kumar Ashutosh, Sarthak Consul, Bhishma Dedhia +3

cs.CV2020

3D-NVS: A 3D Supervision Approach for Next View Selection

Kumar Ashutosh, Saurabh Kumar, Subhasis Chaudhuri

cs.CV2025

FIction: 4D Future Interaction Prediction from Video

Kumar Ashutosh, Georgios Pavlakos, Kristen Grauman

cs.CV2026

SkillSight: Efficient First-Person Skill Assessment with Gaze

Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman

cs.HC2025

Vid2Coach: Transforming How-To Videos into Task Assistants

Mina Huh, Zihui Xue, Ujjaini Das +3

cs.CV2026

Human detectors are surprisingly powerful reward models

Kumar Ashutosh, XuDong Wang, Xi Yin +4

cs.CV2024

SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos

Changan Chen, Kumar Ashutosh, Rohit Girdhar +2

cs.LG2019

Analysis of Lower Bounds for Simple Policy Iteration

Sarthak Consul, Bhishma Dedhia, Kumar Ashutosh +1

cs.CV2025

ExpertAF: Expert Actionable Feedback from Video

Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos +2

cs.CV2026

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions

Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman

cs.CV2024

Learning Object State Changes in Videos: An Open-World Perspective

Zihui Xue, Kumar Ashutosh, Kristen Grauman