6 papers · 1 filter
FoleyBench: A Benchmark For Video-to-Audio Models
Satvik Dixit, Koichi Saito, Zhi Zhong +2
Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects s…
Learning Perceptually Relevant Temporal Envelope Morphing
Satvik Dixit, Sungjoon Park, Chris Donahue +1
Temporal envelope morphing, the process of interpolating between the amplitude dynamics of two audio signals, is an emerging problem in generative audio systems that lacks sufficie…
Mellow: a small audio language model for reasoning
Soham Deshmukh, Satvik Dixit, Rita Singh +1
Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achie…
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
Satvik Dixit, Laurie M. Heller, Chris Donahue
We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct…
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
Satvik Dixit, Soham Deshmukh, Bhiksha Raj
The Automated Audio Captioning (AAC) task aims to describe an audio signal using natural language. To evaluate machine-generated captions, the metrics should take into account audi…
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
Satvik Dixit, Massa Baali, Rita Singh +1
Speaker verification systems have seen significant advancements with the introduction of Multi-scale Feature Aggregation (MFA) architectures, such as MFA-Conformer and ECAPA-TDNN.…