9 papers
The Sharp Dimension Bound in the Johnson--Lindenstrauss Lemma
Vishesh Jain
The Johnson--Lindenstrauss lemma asserts that every set of points in -dimensional Euclidean space embeds into -dimensional Euclidean space with di…
Multimodal Large Language Models as Image Classifiers
Nikita Kisel, Illia Volkov, Klara Janouskova +1
Multimodal Large Language Models (MLLM) classification performance depends critically on evaluation protocol and ground truth quality. Studies comparing MLLMs with supervised and v…
Koo-Fu CLIP: Closed-Form Adaptation of Vision-Language Models via Fukunaga-Koontz Linear Discriminant Analysis
Matej Suchanek, Klara Janouskova, Ondrej Vasatko +1
Visual-language models such as CLIP provide powerful general-purpose representations, but their raw embeddings are not optimized for supervised classification, often exhibiting lim…
Image Recognition with Vision and Language Embeddings of VLMs
Illia Volkov, Nikita Kisel, Klara Janouskova +1
Vision-language models (VLMs) have enabled strong zero-shot classification through image-text alignment. Yet, their purely visual inference capabilities remain under-explored. In t…
SAM2RL: Towards Reinforcement Learning Memory Control in Segment Anything Model 2
Alen Adamyan, Tomáš ÄÞek, Matej Straka +2
Segment Anything Model 2 (SAM 2) has demonstrated strong performance in object segmentation tasks and has become the state-of-the-art for visual object tracking. The model stores i…
FungiTastic: A multi-modal dataset and benchmark for image categorization
Lukas Picek, Klara Janouskova, Vojtech Cermak +1
We introduce a new, challenging benchmark and a dataset, FungiTastic, based on fungal records continuously collected over a twenty-year span. The dataset is labelled and curated by…