3 papers
cs.CV2025
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
Karan Samel, Nitish Sontakke, Irfan Essa
Instructional videos provide a convenient modality to learn new tasks (ex. cooking a recipe, or assembling furniture). A viewer will want to find a corresponding video that reflect…
cs.CV2024
Exploring Efficient Foundational Multi-modal Models for Video Summarization
Karan Samel, Apoorva Beedu, Nitish Sontakke +1
Foundational models are able to generate text outputs given prompt instructions and text, audio, or image inputs. Recently these models have been combined to perform tasks on video…
cs.CV2024
On the Efficacy of Text-Based Input Modalities for Action Anticipation
Apoorva Beedu, Harish Haresamudram, Karan Samel +1
Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down pla…