7 papers
Merlin L48 Spectrogram Dataset
Aaron Sun, Subhransu Maji, Grant Van Horn
In the single-positive multi-label (SPML) setting, each image in a dataset is labeled with the presence of a single class, while the true presence of other classes remains unknown.…
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
Logan Lawrence, Oindrila Saha, Megan Wei +3
Despite the renewed interest in zero-shot visual classification due to the rise of Multimodal Large Language Models (MLLMs), the problem of evaluating free-form responses of auto-r…
Consensus-Driven Active Model Selection
Justin Kay, Grant Van Horn, Subhransu Maji +2
The widespread availability of off-the-shelf machine learning models poses a challenge: which model, of the many available candidates, should be chosen for a given data analysis ta…
Moment Sampling in Video LLMs for Long-Form Video QA
Mustafa Chasmai, Gauri Jagatap, Gouthaman KV +3
Recent advancements in video large language models (Video LLMs) have significantly advanced the field of video question answering (VideoQA). While existing methods perform well on…
Feedforward Few-shot Species Range Estimation
Christian Lange, Max Hamilton, Elijah Cole +6
Knowing where a particular species can or cannot be found on Earth is crucial for ecological research and conservation efforts. By mapping the spatial ranges of all species, we wou…
Generate, Transduct, Adapt: Iterative Transduction with VLMs
Oindrila Saha, Logan Lawrence, Grant Van Horn +1
Transductive zero-shot learning with vision-language models leverages image-image similarities within the dataset to achieve better classification accuracy compared to the inductiv…