Publications (39)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
Lukas Thede, Karsten Roth, Olivier J. Hénaff +2
With the advent and recent ubiquity of foundation models, continual learning (CL) has recently shifted from continual training from scratch to the continual adaptation of pretraine…
Improving the Fairness of Chest X-ray Classifiers
Haoran Zhang, Natalie Dullerud, Karsten Roth +3
Deep learning models have reached or surpassed human-level performance in the field of medical imaging, especially in disease diagnosis using chest x-rays. However, prior work has…
Integrating Language Guidance into Vision-based Deep Metric Learning
Karsten Roth, Oriol Vinyals, Zeynep Akata
Deep Metric Learning (DML) proposes to learn metric spaces which encode semantic similarities as embedding space distances. These spaces should be transferable to classes beyond th…
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
Luca Eyring, Shyamgopal Karthik, Karsten Roth +2
Text-to-Image (T2I) models have made significant advancements in recent years, but they still struggle to accurately capture intricate details specified in complex compositional pr…
ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections
Massimo Bini, Karsten Roth, Zeynep Akata +1
Parameter-efficient finetuning (PEFT) has become ubiquitous to adapt foundation models to downstream task requirements while retaining their generalization ability. However, the am…
COVID-19 Image Data Collection: Prospective Predictions Are the Future
Joseph Paul Cohen, Paul Morrison, Lan Dao +3
Across the world's coronavirus disease 2019 (COVID-19) hot spots, the need to streamline patient diagnosis and management has become more pressing than ever. As one of the main ima…
The Liver Tumor Segmentation Benchmark (LiTS)
Patrick Bilic, Patrick Christ, Hongwei Bran Li +106
In this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedi…
Disentanglement of Correlated Factors via Hausdorff Factorized Support
Karsten Roth, Mark Ibrahim, Zeynep Akata +2
A grand goal in deep learning research is to learn representations capable of generalizing across distribution shifts. Disentanglement is one promising direction aimed at aligning…
A Practitioner's Guide to Continual Multimodal Pretraining
Karsten Roth, Vishaal Udandarao, Sebastian Dziadzio +7
Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over ti…
A Non-isotropic Probabilistic Take on Proxy-based Deep Metric Learning
Michael Kirchhof, Karsten Roth, Zeynep Akata +1
Proxy-based Deep Metric Learning (DML) learns deep representations by embedding images close to their class representatives (proxies), commonly with respect to the angle between th…
PADS: Policy-Adapted Sampling for Visual Similarity Learning
Karsten Roth, Timo Milbich, Björn Ommer
Learning visual similarity requires to learn relations, typically between triplets of images. Albeit triplet approaches being powerful, their computational complexity mostly limits…
If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection
Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini +1
Despite their impressive capabilities, diffusion-based text-to-image (T2I) models can lack faithfulness to the text prompt, where generated images may not contain all the mentioned…
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda +3
Being able to spot defective parts is a critical component in large-scale industrial manufacturing. A particular challenge that we address in this work is the cold-start problem: f…
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
Lukas Thede, Karsten Roth, Matthias Bethge +2
Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but metho…
kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies
Zhongrui Gui, Shuyang Sun, Runjia Li +5
Continual segmentation has not yet tackled the challenge of improving open-vocabulary segmentation models with training data for accurate segmentation across large, continually exp…
Waffling around for Performance: Visual Classification with Random Words and Broad Concepts
Karsten Roth, Jae Myung Kim, A. Sophia Koepke +3
The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as…
Subspace-Boosted Model Merging
Ronald Skorobogat, Karsten Roth, Mariana-Iuliana Georgescu
Model merging enables the combination of multiple specialized expert models into a single model capable of performing multiple tasks. However, the benefits of merging an increasing…
S2SD: Simultaneous Similarity-based Self-Distillation for Deep Metric Learning
Karsten Roth, Timo Milbich, Björn Ommer +2
Deep Metric Learning (DML) provides a crucial tool for visual similarity and zero-shot applications by learning generalizing embedding spaces, although recent work in DML has shown…
Uniform Priors for Data-Efficient Transfer
Samarth Sinha, Karsten Roth, Anirudh Goyal +3
Deep Neural Networks have shown great promise on a variety of downstream applications; but their ability to adapt and generalize to new data and tasks remains a challenge. However,…
Context-Aware Multimodal Pretraining
Karsten Roth, Zeynep Akata, Dima Damen +2
Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amou…
Is Fairness Only Metric Deep? Evaluating and Addressing Subgroup Gaps in Deep Metric Learning
Natalie Dullerud, Karsten Roth, Kimia Hamidieh +2
Deep metric learning (DML) enables learning with less supervision through its emphasis on the similarity structure of representations. There has been much work on improving general…
Characterizing Generalization under Out-Of-Distribution Shifts in Deep Metric Learning
Timo Milbich, Karsten Roth, Samarth Sinha +3
Deep Metric Learning (DML) aims to find representations suitable for zero-shot transfer to a priori unknown test distributions. However, common evaluation protocols only test a sin…
Predicting COVID-19 Pneumonia Severity on Chest X-ray with Deep Learning
Joseph Paul Cohen, Lan Dao, Paul Morrison +8
Purpose: The need to streamline patient management for COVID-19 has become more pressing than ever. Chest X-rays provide a non-invasive (potentially bedside) tool to monitor the pr…
Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning
Zafir Stojanovski, Karsten Roth, Zeynep Akata
Large pre-trained, zero-shot capable models have shown considerable success both for standard transfer and adaptation tasks, with particular robustness towards distribution shifts.…
MIC: Mining Interclass Characteristics for Improved Metric Learning
Karsten Roth, Biagio Brattoli, Björn Ommer
Metric learning seeks to embed images of objects suchthat class-defined relations are captured by the embeddingspace. However, variability in images is not just due to different de…
DiVA: Diverse Visual Feature Aggregation for Deep Metric Learning
Timo Milbich, Karsten Roth, Homanga Bharadhwaj +4
Visual Similarity plays an important role in many computer vision applications. Deep metric learning (DML) is a powerful framework for learning such similarities which not only gen…
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
Nishad Singhi, Jae Myung Kim, Karsten Roth +1
Concept Bottleneck Models (CBMs) ground image classification on human-understandable concepts to allow for interpretable model decisions. Crucially, the CBM design inherently allow…
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Matteo Farina, Vishaal Udandarao, Thao Nguyen +34
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curat…
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model
Karsten Roth, Lukas Thede, Almut Sophia Koepke +3
Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization. In this work, we find these training variati…
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…
Revisiting Training Strategies and Generalization Performance in Deep Metric Learning
Karsten Roth, Timo Milbich, Samarth Sinha +3
Deep Metric Learning (DML) is arguably one of the most influential lines of research for learning visual similarities with many proposed approaches every year. Although the field b…
Vision-by-Language for Training-Free Compositional Image Retrieval
Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini +1
Given an image and a target modification (e.g an image of the Eiffel tower and the text "without people and at night-time"), Compositional Image Retrieval (CIR) aims to retrieve th…
How to Merge Your Multimodal Models Over Time?
Sebastian Dziadzio, Vishaal Udandarao, Karsten Roth +4
Model merging combines multiple expert models - finetuned from a base foundation model on diverse tasks and domains - into a single, more capable model. However, most existing mode…
Mask Mining for Improved Liver Lesion Segmentation
Karsten Roth, Jürgen Hesser, Tomasz KonopczyÅski
We propose a novel procedure to improve liver and lesion segmentation from CT scans for U-Net based models. Our method extends standard segmentation pipelines to focus on higher ta…
Liver Lesion Segmentation with slice-wise 2D Tiramisu and Tversky loss function
Karsten Roth, Tomasz KonopczyÅski, Jürgen Hesser
At present, lesion segmentation is still performed manually (or semi-automatically) by medical experts. To facilitate this process, we contribute a fully-automatic lesion segmentat…
Disentangled Representation Learning with the Gromov-Monge Gap
Théo Uscidda, Luca Eyring, Karsten Roth +3
Learning disentangled representations from unlabelled data is a fundamental challenge in machine learning. Solving it may unlock other problems, such as generalization, interpretab…
Sharing Matters for Generalization in Deep Metric Learning
Timo Milbich, Karsten Roth, Biagio Brattoli +1
Learning the similarity between images constitutes the foundation for numerous vision tasks. The common paradigm is discriminative metric learning, which seeks an embedding that se…
Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
Yiran Huang, Karsten Roth, Quentin Bouniot +2
Transformer-based multimodal large language models often exhibit in-context learning (ICL) abilities. Motivated by this phenomenon, we ask: how do transformers learn to associate i…
Non-isotropy Regularization for Proxy-based Deep Metric Learning
Karsten Roth, Oriol Vinyals, Zeynep Akata
Deep Metric Learning (DML) aims to learn representation spaces on which semantic relations can simply be expressed through predefined distance metrics. Best performing approaches c…