12 papers
Bioacoustic Geolocation: Species Sounds as Geographic Signals
Mustafa Chasmai, Wuao Liu, Subhransu Maji +1
Can we determine someone's geographic location solely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? In this work, we tac…
Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
Wuao Liu, Mustafa Chasmai, Subhransu Maji +1
Bioacoustic recognition requires fine-grained acoustic understanding to distinguish similar-sounding species. However, many large-scale data repositories such as iNaturalist are we…
RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
Logan Lawrence, Mustafa Chasmai, Rangel Daroya +8
Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, c…
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
Logan Lawrence, Oindrila Saha, Megan Wei +3
Despite the renewed interest in zero-shot visual classification due to the rise of Multimodal Large Language Models (MLLMs), the problem of evaluating free-form responses of auto-r…
Merlin L48 Spectrogram Dataset
Aaron Sun, Subhransu Maji, Grant Van Horn
In the single-positive multi-label (SPML) setting, each image in a dataset is labeled with the presence of a single class, while the true presence of other classes remains unknown.…
Generate, Transduct, Adapt: Iterative Transduction with VLMs
Oindrila Saha, Logan Lawrence, Grant Van Horn +1
Transductive zero-shot learning with vision-language models leverages image-image similarities within the dataset to achieve better classification accuracy compared to the inductiv…