8 papers
Event-Grounded Question Answering over Long Audio via Structured Retrieval
Kartik Hegde, Arvind Krishna Sridhar, Naveen Vakada +2
Answering natural-language questions over multi-hour audio requires reliable event recognition, temporal grounding, and efficient retrieval. We present LA-RAG (Long Audio Retrieval…
Aligning Audio Captions with Human Preferences
Kartik Hegde, Rehana Mahfuz, Yinyi Guo +1
Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To a…
Proactive Conversational Assistant for a Procedural Manual Task based on Audio and IMU
Rehana Mahfuz, Yinyi Guo, Erik Visser +1
Real-time conversational assistants for procedural manual tasks often depend on video input, which can be computationally expensive and compromise user privacy. For the first time,…
Spatial Audio Question Answering and Reasoning on Dynamic Source Movements
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser
Spatial audio understanding aims to enable machines to interpret complex auditory scenes, particularly when sound sources move over time. In this work, we study Spatial Audio Quest…
Spatial Audio Motion Understanding and Reasoning
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser
Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding wi…
Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser
The Audio Question Answering (AQA) task includes audio event classification, audio captioning, and open-ended reasoning. Recently, AQA has garnered attention due to the advent of L…