15 citations · 20 across the 5 of their papers we have counts for
3 papers · 1 filter
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
Karan Samel, Nitish Sontakke, Irfan Essa
Instructional videos provide a convenient modality to learn new tasks (ex. cooking a recipe, or assembling furniture). A viewer will want to find a corresponding video that reflect…
Exploring Efficient Foundational Multi-modal Models for Video Summarization
Karan Samel, Apoorva Beedu, Nitish Sontakke +1
Foundational models are able to generate text outputs given prompt instructions and text, audio, or image inputs. Recently these models have been combined to perform tasks on video…
On the Efficacy of Text-Based Input Modalities for Action Anticipation
Apoorva Beedu, Harish Haresamudram, Karan Samel +1
Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down pla…