activity
20192025
most citedSpeechVerse: A Large-scale Generalizable Audio Language Model

5 citations · 7 across the 5 of their papers we have counts for

collaborators

10 papers

cs.HC2025

TILES-2018 Sleep Benchmark Dataset: A Longitudinal Wearable Sleep Data Set of Hospital Workers for Modeling and Understanding Sleep Behaviors

Tiantian Feng, Brandon M Booth, Karel Mundnich +4

Sleep is important for everyday functioning, overall well-being, and quality of life. Recent advances in wearable sensing technology have enabled continuous, noninvasive, and cost-…

eess.AS2025

Speech Retrieval-Augmented Generation without Automatic Speech Recognition

Do June Min, Karel Mundnich, Andy Lapastora +3

One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented ge…

cs.CL2024

Zero-resource Speech Translation and Recognition with LLMs

Karel Mundnich, Xing Niu, Prashant Mathur +10

Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose…

cs.CL2024

SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models

Raghuveer Peri, Sai Muralidhar Jayanthi, Srikanth Ronanki +11

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and r…

cs.CL2024★ 5 cited

SpeechVerse: A Large-scale Generalizable Audio Language Model

Nilaksh Das, Saket Dingliwal, Srikanth Ronanki +14

Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have f…

cs.CV2021

Audiovisual Highlight Detection in Videos

Karel Mundnich, Alexandra Fenster, Aparna Khare +1

In this paper, we test the hypothesis that interesting events in unstructured videos are inherently audiovisual. We combine deep image representations for object recognition and sc…