86 citations · 181 across the 35 of their papers we have counts for
15 papers · 1 filter
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker +4
Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace t…
SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability
Amin Karimi Monsefi, Abolfazl Meyarian, Mridul Khurana +6
Animals are described as effectively camouflaged when they blend seamlessly with their surrounding, yet no standardized quantitative measure of this seamlessness exists. We address…
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life
Mridul Khurana, Amin Karimi Monsefi, Justin Lee +9
Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the rema…
A continental-scale dataset of ground beetles with high-resolution images and validated morphological trait measurements
S M Rayeed, Mridul Khurana, Alyson East +18
Despite the ecological significance of invertebrates, global trait databases remain heavily biased toward vertebrates and plants, limiting comprehensive ecological analyses of high…
TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation
Amin Karimi Monsefi, Mridul Khurana, Rajiv Ramnath +3
We propose TaxaDiffusion, a taxonomy-informed training framework for diffusion models to generate fine-grained animal images with high morphological and identity accuracy. Unlike s…
Open World Scene Graph Generation using Vision Language Models
Amartya Dutta, Kazi Sajeed Mehrab, Medha Sawhney +8
Scene-Graph Generation (SGG) seeks to recognize objects in an image and distill their salient pairwise relationships. Most methods depend on dataset-specific supervision to learn t…