papers

Publications (39)

cs.LG2025

Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models

Lukas Thede, Karsten Roth, Olivier J. Hénaff +2

With the advent and recent ubiquity of foundation models, continual learning (CL) has recently shifted from continual training from scratch to the continual adaptation of pretraine…

cs.LG2022

Improving the Fairness of Chest X-ray Classifiers

Haoran Zhang, Natalie Dullerud, Karsten Roth +3

Deep learning models have reached or surpassed human-level performance in the field of medical imaging, especially in disease diagnosis using chest x-rays. However, prior work has…

cs.CV2022

Integrating Language Guidance into Vision-based Deep Metric Learning

Karsten Roth, Oriol Vinyals, Zeynep Akata

Deep Metric Learning (DML) proposes to learn metric spaces which encode semantic similarities as embedding space distances. These spaces should be transferable to classes beyond th…

cs.CV2024

ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization

Luca Eyring, Shyamgopal Karthik, Karsten Roth +2

Text-to-Image (T2I) models have made significant advancements in recent years, but they still struggle to accurately capture intricate details specified in complex compositional pr…

cs.LG2024

ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections

Massimo Bini, Karsten Roth, Zeynep Akata +1

Parameter-efficient finetuning (PEFT) has become ubiquitous to adapt foundation models to downstream task requirements while retaining their generalization ability. However, the am…

q-bio.QM2020

COVID-19 Image Data Collection: Prospective Predictions Are the Future

Joseph Paul Cohen, Paul Morrison, Lan Dao +3

Across the world's coronavirus disease 2019 (COVID-19) hot spots, the need to streamline patient diagnosis and management has become more pressing than ever. As one of the main ima…

cs.CV2022

The Liver Tumor Segmentation Benchmark (LiTS)

Patrick Bilic, Patrick Christ, Hongwei Bran Li +106

In this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedi…

cs.LG2023

Disentanglement of Correlated Factors via Hausdorff Factorized Support

Karsten Roth, Mark Ibrahim, Zeynep Akata +2

A grand goal in deep learning research is to learn representations capable of generalizing across distribution shifts. Disentanglement is one promising direction aimed at aligning…

cs.CV2024

A Practitioner's Guide to Continual Multimodal Pretraining

Karsten Roth, Vishaal Udandarao, Sebastian Dziadzio +7

Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over ti…

cs.LG2022

A Non-isotropic Probabilistic Take on Proxy-based Deep Metric Learning

Michael Kirchhof, Karsten Roth, Zeynep Akata +1

Proxy-based Deep Metric Learning (DML) learns deep representations by embedding images close to their class representatives (proxies), commonly with respect to the angle between th…

cs.CV2020

PADS: Policy-Adapted Sampling for Visual Similarity Learning

Karsten Roth, Timo Milbich, Björn Ommer

Learning visual similarity requires to learn relations, typically between triplets of images. Albeit triplet approaches being powerful, their computational complexity mostly limits…

cs.CV2023

If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection

Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini +1

Despite their impressive capabilities, diffusion-based text-to-image (T2I) models can lack faithfulness to the text prompt, where generated images may not contain all the mentioned…

cs.CV2022

Towards Total Recall in Industrial Anomaly Detection

Karsten Roth, Latha Pemula, Joaquin Zepeda +3

Being able to spot defective parts is a critical component in large-scale industrial manufacturing. A particular challenge that we address in this work is the cold-start problem: f…

cs.CL2025

WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs

Lukas Thede, Karsten Roth, Matthias Bethge +2

Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but metho…

cs.CV2024

kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies

Zhongrui Gui, Shuyang Sun, Runjia Li +5

Continual segmentation has not yet tackled the challenge of improving open-vocabulary segmentation models with training data for accurate segmentation across large, continually exp…

cs.CV2023

Waffling around for Performance: Visual Classification with Random Words and Broad Concepts

Karsten Roth, Jae Myung Kim, A. Sophia Koepke +3

The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as…

cs.LG2025

Subspace-Boosted Model Merging

Ronald Skorobogat, Karsten Roth, Mariana-Iuliana Georgescu

Model merging enables the combination of multiple specialized expert models into a single model capable of performing multiple tasks. However, the benefits of merging an increasing…

cs.CV2021

S2SD: Simultaneous Similarity-based Self-Distillation for Deep Metric Learning

Karsten Roth, Timo Milbich, Björn Ommer +2

Deep Metric Learning (DML) provides a crucial tool for visual similarity and zero-shot applications by learning generalizing embedding spaces, although recent work in DML has shown…

cs.LG2020

Uniform Priors for Data-Efficient Transfer

Samarth Sinha, Karsten Roth, Anirudh Goyal +3

Deep Neural Networks have shown great promise on a variety of downstream applications; but their ability to adapt and generalize to new data and tasks remains a challenge. However,…

cs.CV2024

Context-Aware Multimodal Pretraining

Karsten Roth, Zeynep Akata, Dima Damen +2

Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amou…

cs.LG2022

Is Fairness Only Metric Deep? Evaluating and Addressing Subgroup Gaps in Deep Metric Learning

Natalie Dullerud, Karsten Roth, Kimia Hamidieh +2

Deep metric learning (DML) enables learning with less supervision through its emphasis on the similarity structure of representations. There has been much work on improving general…

cs.LG2021

Characterizing Generalization under Out-Of-Distribution Shifts in Deep Metric Learning

Timo Milbich, Karsten Roth, Samarth Sinha +3

Deep Metric Learning (DML) aims to find representations suitable for zero-shot transfer to a priori unknown test distributions. However, common evaluation protocols only test a sin…

eess.IV2020

Predicting COVID-19 Pneumonia Severity on Chest X-ray with Deep Learning

Joseph Paul Cohen, Lan Dao, Paul Morrison +8

Purpose: The need to streamline patient management for COVID-19 has become more pressing than ever. Chest X-rays provide a non-invasive (potentially bedside) tool to monitor the pr…

cs.LG2022

Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning

Zafir Stojanovski, Karsten Roth, Zeynep Akata

Large pre-trained, zero-shot capable models have shown considerable success both for standard transfer and adaptation tasks, with particular robustness towards distribution shifts.…

cs.CV2019

MIC: Mining Interclass Characteristics for Improved Metric Learning

Karsten Roth, Biagio Brattoli, Björn Ommer

Metric learning seeks to embed images of objects suchthat class-defined relations are captured by the embeddingspace. However, variability in images is not just due to different de…

cs.CV2020

DiVA: Diverse Visual Feature Aggregation for Deep Metric Learning

Timo Milbich, Karsten Roth, Homanga Bharadhwaj +4

Visual Similarity plays an important role in many computer vision applications. Deep metric learning (DML) is a powerful framework for learning such similarities which not only gen…

cs.LG2024

Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models

Nishad Singhi, Jae Myung Kim, Karsten Roth +1

Concept Bottleneck Models (CBMs) ground image classification on human-understandable concepts to allow for interpretable model decisions. Crucially, the CBM design inherently allow…

cs.CV2026

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Matteo Farina, Vishaal Udandarao, Thao Nguyen +34

Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curat…

cs.LG2024

Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model

Karsten Roth, Lukas Thede, Almut Sophia Koepke +3

Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization. In this work, we find these training variati…

cs.CL2026

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…

cs.CV2020

Revisiting Training Strategies and Generalization Performance in Deep Metric Learning

Karsten Roth, Timo Milbich, Samarth Sinha +3

Deep Metric Learning (DML) is arguably one of the most influential lines of research for learning visual similarities with many proposed approaches every year. Although the field b…

cs.CV2024

Vision-by-Language for Training-Free Compositional Image Retrieval

Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini +1

Given an image and a target modification (e.g an image of the Eiffel tower and the text "without people and at night-time"), Compositional Image Retrieval (CIR) aims to retrieve th…

cs.LG2024

How to Merge Your Multimodal Models Over Time?

Sebastian Dziadzio, Vishaal Udandarao, Karsten Roth +4

Model merging combines multiple expert models - finetuned from a base foundation model on diverse tasks and domains - into a single, more capable model. However, most existing mode…

eess.IV2020

Mask Mining for Improved Liver Lesion Segmentation

Karsten Roth, Jürgen Hesser, Tomasz Konopczyński

We propose a novel procedure to improve liver and lesion segmentation from CT scans for U-Net based models. Our method extends standard segmentation pipelines to focus on higher ta…

cs.CV2019

Liver Lesion Segmentation with slice-wise 2D Tiramisu and Tversky loss function

Karsten Roth, Tomasz Konopczyński, Jürgen Hesser

At present, lesion segmentation is still performed manually (or semi-automatically) by medical experts. To facilitate this process, we contribute a fully-automatic lesion segmentat…

cs.LG2025

Disentangled Representation Learning with the Gromov-Monge Gap

Théo Uscidda, Luca Eyring, Karsten Roth +3

Learning disentangled representations from unlabelled data is a fundamental challenge in machine learning. Solving it may unlock other problems, such as generalization, interpretab…

cs.CV2021

Sharing Matters for Generalization in Deep Metric Learning

Timo Milbich, Karsten Roth, Biagio Brattoli +1

Learning the similarity between images constitutes the foundation for numerous vision tasks. The common paradigm is discriminative metric learning, which seeks an embedding that se…

cs.CL2026

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

Yiran Huang, Karsten Roth, Quentin Bouniot +2

Transformer-based multimodal large language models often exhibit in-context learning (ICL) abilities. Motivated by this phenomenon, we ask: how do transformers learn to associate i…

cs.CV2022

Non-isotropy Regularization for Proxy-based Deep Metric Learning

Karsten Roth, Oriol Vinyals, Zeynep Akata

Deep Metric Learning (DML) aims to learn representation spaces on which semantic relations can simply be expressed through predefined distance metrics. Best performing approaches c…