6 papers
The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models
Robert Welch, Emir Konuk, Kevin Smith
Vision-language models (VLMs) are increasingly deployed in high-stakes settings where reliable uncertainty quantification (UQ) is as important as predictive accuracy. Extended reas…
Learning What Helps: Task-Aligned Context Selection for Vision Tasks
Jingyu Guo, Emir Konuk, Fredrik Strand +2
Humans often resolve visual uncertainty by comparing an image with relevant examples, but ViTs lack the ability to identify which examples would improve their predictions. We prese…
Efficient Self-Supervised Adaptation for Medical Image Analysis
Moein Sorkhei, Emir Konuk, Jingyu Guo +3
Self-supervised adaptation (SSA) improves foundation model transfer to medical domains but is computationally prohibitive. Although parameter efficient fine-tuning methods such as…
APLA: A Simple Adaptation Method for Vision Transformers
Moein Sorkhei, Emir Konuk, Kevin Smith +1
Existing adaptation techniques typically require architectural modifications or added parameters, leading to high computational costs and complexity. We introduce Attention Project…
VORTEX: Challenging CNNs at Texture Recognition by using Vision Transformers with Orderless and Randomized Token Encodings
Leonardo Scabini, Kallil M. Zielinski, Emir Konuk +4
Texture recognition has recently been dominated by ImageNet-pre-trained deep Convolutional Neural Networks (CNNs), with specialized modifications and feature engineering required t…
Learning from Offline Foundation Features with Tensor Augmentations
Emir Konuk, Christos Matsoukas, Moein Sorkhei +2
We introduce Learning from Offline Foundation Features with Tensor Augmentations (LOFF-TA), an efficient training scheme designed to harness the capabilities of foundation models i…