3 citations · 8 across the 8 of their papers we have counts for
9 papers
Proactive Conversational Assistant for a Procedural Manual Task based on Audio and IMU
Rehana Mahfuz, Yinyi Guo, Erik Visser +1
Real-time conversational assistants for procedural manual tasks often depend on video input, which can be computationally expensive and compromise user privacy. For the first time,…
Aligning Audio Captions with Human Preferences
Kartik Hegde, Rehana Mahfuz, Yinyi Guo +1
Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To a…
Resource-Efficient Reference-Free Evaluation of Audio Captions
Rehana Mahfuz, Yinyi Guo, Erik Visser
To establish the trustworthiness of systems that automatically generate text captions for audio, images and video, existing reference-free metrics rely on large pretrained models w…
Parameter Efficient Audio Captioning With Faithful Guidance Using Audio-text Shared Latent Representation
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser +1
There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models are fre…
Detecting False Alarms and Misses in Audio Captions
Rehana Mahfuz, Yinyi Guo, Arvind Krishna Sridhar +1
Metrics to evaluate audio captions simply provide a score without much explanation regarding what may be wrong in case the score is low. Manual human intervention is needed to find…
Mitigating Gradient-based Adversarial Attacks via Denoising and Compression
Rehana Mahfuz, Rajeev Sahay, Aly El Gamal
Gradient-based adversarial attacks on deep neural networks pose a serious threat, since they can be deployed by adding imperceptible perturbations to the test data of any network,…