activity
20242026
most citedGemma 4 Technical Report

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

Efficient Randomized Experiments Using Foundation Models

Piersilvio De Bartolomeis, Javier Abad, Guanbo Wang +4

Randomized experiments are the preferred approach for evaluating the effects of interventions, but they are costly and often yield estimates with substantial uncertainty. On the ot…

cs.LG2025

Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation models

Konstantin Donhauser, Kristina Ulicna, Gemma Elyse Moran +4

Sparse dictionary learning (DL) has emerged as a powerful approach to extract semantically meaningful concepts from the internals of large language models (LLMs) trained mainly in…

cs.LG2025

ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy

Kian Kenyon-Dean, Zitong Jerry Wang, John Urbanik +10

Large-scale cell microscopy screens are used in drug discovery and molecular biology research to study the effects of millions of chemical and genetic perturbations on cells. To us…

cs.LG2024

Copyright-Protected Language Generation via Adaptive Model Fusion

Javier Abad, Konstantin Donhauser, Francesco Pinto +1

The risk of language models reproducing copyrighted material from their training data has led to the development of various protective measures. Among these, inference-time strateg…

cs.LG2024

Strong Copyright Protection for Language Models via Adaptive Model Fusion

Javier Abad, Konstantin Donhauser, Francesco Pinto +1

The risk of language models unintentionally reproducing copyrighted material from their training data has led to the development of various protective measures. In this paper, we p…

cs.LG2024

Privacy-preserving data release leveraging optimal transport and particle gradient descent

Konstantin Donhauser, Javier Abad, Neha Hulkund +1

We present a novel approach for differentially private data synthesis of protected tabular datasets, a relevant task in highly sensitive domains such as healthcare and government.…