3 papers
eess.AS2024
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding
Vakada Naveen, Arvind Krishna Sridhar, Yinyi Guo +1
This paper presents a comprehensive chatbot system designed to handle a wide range of audio-related queries by integrating multiple specialized audio processing models. The propose…
cs.CL2023
Parameter Efficient Audio Captioning With Faithful Guidance Using Audio-text Shared Latent Representation
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser +1
There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models are fre…
cs.MM2023
Detecting False Alarms and Misses in Audio Captions
Rehana Mahfuz, Yinyi Guo, Arvind Krishna Sridhar +1
Metrics to evaluate audio captions simply provide a score without much explanation regarding what may be wrong in case the score is low. Manual human intervention is needed to find…