1 paper
Satvik Dixit, Laurie M. Heller, Chris Donahue
We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct…