4 papers
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Ramit Pahwa, Apoorva Beedu, Parivesh Priye +4
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acousti…
HierSum: A Global and Local Attention Mechanism for Video Summarization
Apoorva Beedu, Irfan Essa
Video summarization creates an abridged version (i.e., a summary) that provides a quick overview of the video while retaining pertinent information. In this work, we focus on summa…
Mamba Fusion: Learning Actions Through Questioning
Zhikang Dong, Apoorva Beedu, Jason Sheinkopf +1
Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de f…
Exploring Efficient Foundational Multi-modal Models for Video Summarization
Karan Samel, Apoorva Beedu, Nitish Sontakke +1
Foundational models are able to generate text outputs given prompt instructions and text, audio, or image inputs. Recently these models have been combined to perform tasks on video…