20 citations · 41 across the 21 of their papers we have counts for
3 papers · 2 filters
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a sin…
Mellow: a small audio language model for reasoning
Soham Deshmukh, Satvik Dixit, Rita Singh +1
Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achie…
ADIFF: Explaining audio difference using natural language
Soham Deshmukh, Shuo Han, Rita Singh +1
Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying an…