paper

Multimodal Taxonomic Conditioning for Generative Plankton Imagery

arXiv:2609.11673

Abstract

Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer. We evaluate synthetic sample quality on distributional fidelity and downstream classifier utility.

European Conference on Computer Vision (ECCV) 2nd Workshop on Marine Vision

Multimodal Taxonomic Conditioning for Generative Plankton Imagery · wovepaper