papers

Publications (27)

cs.LG2025

Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition

Ilker Demirel, Karan Thakkar, Benjamin Elizalde +7

Sensor data streams provide valuable information around activities and context for downstream applications, though integrating complementary information can be challenging. We show…

cs.SD2024

Natural Language Supervision for General-Purpose Audio Representations

Benjamin Elizalde, Soham Deshmukh, Huaming Wang

Audio-Language models jointly learn multimodal text and audio representations that enable Zero-Shot inference. Models rely on the encoders to create powerful representations of the…

cs.SD2018

Framework for evaluation of sound event detection in web videos

Rohan Badlani, Ankit Shah, Benjamin Elizalde +2

The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a mat…

cs.SD2018

Content-based Representations of audio using Siamese neural networks

Pranay Manocha, Rohan Badlani, Anurag Kumar +3

In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This pro…

cs.SD2016

Experiments on the DCASE Challenge 2016: Acoustic Scene Classification and Sound Event Detection in Real Life Recording

Benjamin Elizalde, Anurag Kumar, Ankit Shah +4

In this paper we present our work on Task 1 Acoustic Scene Classi- fication and Task 3 Sound Event Detection in Real Life Recordings. Among our experiments we have low-level and hi…

cs.SD2023

Prompting Audios Using Acoustic Properties For Emotion Representation

Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh +3

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion…

cs.MM2015

The YLI-MED Corpus: Characteristics, Procedures, and Plans

Julia Bernd, Damian Borth, Benjamin Elizalde +7

The YLI Multimedia Event Detection corpus is a public-domain index of videos with annotations and computed features, specialized for research in multimedia event detection (MED), i…

cs.SD2018

AudioPairBank: Towards A Large-Scale Tag-Pair-Based Audio Content Analysis

Sebastian Sager, Benjamin Elizalde, Damian Borth +3

Recently, sound recognition has been used to identify sounds, such as car and river. However, sounds have nuances that may be better described by adjective-noun pairs such as slow…

cs.SD2023

NELS -- Never-Ending Learner of Sounds

Benjamin Elizalde, Rohan Badlani, Ankit Shah +2

Sounds are essential to how humans perceive and interact with the world and are captured in recordings and shared on the Internet on a minute-by-minute basis. These recordings, whi…

eess.AS2023

Training Audio Captioning Models without Audio

Soham Deshmukh, Benjamin Elizalde, Dimitra Emmanouilidou +3

Automated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audi…

eess.AS2023

Synergy between human and machine approaches to sound/scene recognition and processing: An overview of ICASSP special session

Laurie M. Heller, Benjamin Elizalde, Bhiksha Raj +1

Machine Listening, as usually formalized, attempts to perform a task that is, from our perspective, fundamentally human-performable, and performed by humans. Current automated mode…

eess.AS2020

Multi-label Sound Event Retrieval Using a Deep Learning-based Siamese Structure with a Pairwise Presence Matrix

Jianyu Fan, Eric Nichols, Daniel Tompkins +3

Realistic recordings of soundscapes often have multiple sound events co-occurring, such as car horns, engine and human voices. Sound event retrieval is a type of content-based sear…

cs.SD2021

Identifying Actions for Sound Event Classification

Benjamin Elizalde, Radu Revutchi, Samarjit Das +3

In Psychology, actions are paramount for humans to identify sound events. In Machine Learning (ML), action recognition achieves high accuracy; however, it has not been asked whethe…

cs.SD2022

Describing emotions with acoustic property prompts for speech emotion recognition

Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh +3

Emotions lie on a broad continuum and treating emotions as a discrete number of classes limits the ability of a model to capture the nuances in the continuum. The challenge is how…

cs.SD2016

Audio Content based Geotagging in Multimedia

Anurag Kumar, Benjamin Elizalde, Bhiksha Raj

In this paper we propose methods to extract geographically relevant information in a multimedia recording using its audio. Our method primarily is based on the fact that urban acou…

eess.AS2017

Audio Concept Classification with Hierarchical Deep Neural Networks

Mirco Ravanelli, Benjamin Elizalde, Karl Ni +1

Audio-based multimedia retrieval tasks may identify semantic information in audio streams, i.e., audio concepts (such as music, laughter, or a revving engine). Conventional Gaussia…

cs.SD2025

Discrete Audio Tokens: More Than a Survey!

Pooneh Mousavi, Gallil Maimon, Adel Moumen +18

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and infere…

cs.SD2021

COVID-19 Detection Using Recorded Coughs in the 2021 DiCOVA Challenge

Benjamin Elizalde, Daniel Tompkins

COVID-19 has resulted in over 100 million infections and caused worldwide lock downs due to its high transmission rate and limited testing options. Current diagnostic tests can be…

cs.SD2018

DCASE 2017 Task 1: Acoustic Scene Classification Using Shift-Invariant Kernels and Random Features

Abelino Jimenez, Benjamin Elizalde, Bhiksha Raj

Acoustic scene recordings are represented by different types of handcrafted or Neural Network-derived features. These features, typically of thousands of dimensions, are classified…

eess.AS2024

PAM: Prompting Audio-Language Models for Audio Quality Assessment

Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde +5

While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge. Audio-Language Mod…

eess.AS2022

Audio Retrieval with WavText5K and CLAP Training

Soham Deshmukh, Benjamin Elizalde, Huaming Wang

Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve rele…

cs.MM2016

City-Identification of Flickr Videos Using Semantic Acoustic Features

Benjamin Elizalde, Guan-Lin Chao, Ming Zeng +1

City-identification of videos aims to determine the likelihood of a video belonging to a set of cities. In this paper, we present an approach using only audio, thus we do not use a…

cs.SD2024

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

Soham Deshmukh, Shuo Han, Hazim Bukhari +4

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable perfor…

eess.AS2024

Pengi: An Audio Language Model for Audio Tasks

Soham Deshmukh, Benjamin Elizalde, Rita Singh +1

In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the develo…

cs.SD2017

An Approach for Self-Training Audio Event Detectors Using Web Data

Benjamin Elizalde, Ankit Shah, Siddharth Dalmia +5

Audio Event Detection (AED) aims to recognize sounds within audio and video recordings. AED employs machine learning algorithms commonly trained and tested on annotated datasets. H…

cs.SD2022

CLAP: Learning Audio Concepts From Natural Language Supervision

Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail +1

Mainstream Audio Analytics models are trained to learn under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision lim…

cs.MM2016

YFCC100M: The New Data in Multimedia Research

Bart Thomee, David A. Shamma, Gerald Friedland +5

We present the Yahoo Flickr Creative Commons 100 Million Dataset (YFCC100M), the largest public multimedia collection that has ever been released. The dataset contains a total of 1…