works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2026

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4

OpenBEATs is an open-source framework that extends the BEATs audio encoder with multi-domain masked-token pretraining, achieving state-of-the-art results on a wide range of audio t…

cs.SD2026

The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge

Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4

This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…

cs.SD2025

Mellow: a small audio language model for reasoning

Soham Deshmukh, Satvik Dixit, Rita Singh +1

Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achie…

cs.SD2025

ADIFF: Explaining audio difference using natural language

Soham Deshmukh, Shuo Han, Rita Singh +1

Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying an…

cs.SD2024

MACE: Leveraging Audio for Evaluating Audio Captioning Systems

Satvik Dixit, Soham Deshmukh, Bhiksha Raj

The Automated Audio Captioning (AAC) task aims to describe an audio signal using natural language. To evaluate machine-generated captions, the metrics should take into account audi…

cs.SD2024

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

Soham Deshmukh, Shuo Han, Hazim Bukhari +4

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable perfor…