14 papers
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson +9
Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing…
Cross-Modal Corroboration for Annotation-Free Wildlife Monitoring
Bharath Pillai, Varun Viswapriyan, Christopher Stewart +2
Scaling wildlife monitoring for real-world conservation deployments requires automated analysis of smart sensors that operate under severe annotation scarcity. We propose leveragin…
Leveraging Latent Visual Reasoning in Silence
Dongyao Zhu, Zhen Wang, Xi Xiao +7
Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
Ziheng Zhang, Xinyue Ma, Arpita Chowdhury +9
This work investigates descriptive captions as an additional source of supervision for biological multimodal foundation models. Images and captions can be viewed as complementary s…
A continental-scale dataset of ground beetles with high-resolution images and validated morphological trait measurements
S M Rayeed, Mridul Khurana, Alyson East +18
Despite the ecological significance of invertebrates, global trait databases remain heavily biased toward vertebrates and plants, limiting comprehensive ecological analyses of high…
Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf +1
Scientific archives now contain hundreds of petabytes of data across genomics, ecology, climate, and molecular biology that could reveal undiscovered patterns if systematically ana…