From the 1 of 8 linked papers with an AI index.
8 papers · 1 filter
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
OpenBEATs is an open-source framework that extends the BEATs audio encoder with multi-domain masked-token pretraining, achieving state-of-the-art results on a wide range of audio t…
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…
Mellow: a small audio language model for reasoning
Soham Deshmukh, Satvik Dixit, Rita Singh +1
Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achie…
ADIFF: Explaining audio difference using natural language
Soham Deshmukh, Shuo Han, Rita Singh +1
Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying an…
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
Satvik Dixit, Soham Deshmukh, Bhiksha Raj
The Automated Audio Captioning (AAC) task aims to describe an audio signal using natural language. To evaluate machine-generated captions, the metrics should take into account audi…
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
Soham Deshmukh, Shuo Han, Hazim Bukhari +4
Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable perfor…