16 papers
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker +4
Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace t…
LakeFM: Toward a Foundation Model for Aquatic Ecosystems Using Irregular Multivariate Multi-depth Time Series Data
Abhilash Neog, Sepideh Fatemi, Medha Sawhney +9
Understanding and forecasting lake dynamics is critical for monitoring water quality and ecosystem health across lakes and reservoirs. While machine learning methods have been rece…
SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability
Amin Karimi Monsefi, Abolfazl Meyarian, Mridul Khurana +6
Animals are described as effectively camouflaged when they blend seamlessly with their surrounding, yet no standardized quantitative measure of this seamlessness exists. We address…
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life
Mridul Khurana, Amin Karimi Monsefi, Justin Lee +9
Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the rema…
VILLA: Versatile Information Retrieval From Scientific Literature Using Large LAnguage Models
Blessy Antony, Amartya Dutta, Sneha Aggarwal +7
The lack of high-quality ground truth datasets to train machine learning (ML) models impedes the potential of artificial intelligence (AI) for science research. Scientific informat…
A continental-scale dataset of ground beetles with high-resolution images and validated morphological trait measurements
S M Rayeed, Mridul Khurana, Alyson East +18
Despite the ecological significance of invertebrates, global trait databases remain heavily biased toward vertebrates and plants, limiting comprehensive ecological analyses of high…