collaborators

6 papers

cs.CL2025

Scaling Embedding Layers in Language Models

Da Yu, Edith Cohen, Badih Ghazi +5

We propose (calable, ontextualized, ffloaded, -gram mbedding), a new method for extending input embedding layers to enhance language model performance. To av…

cs.CR2025

VaultGemma: A Differentially Private Gemma Model

Amer Sinha, Thomas Mesnard, Ryan McKenna +18

We introduce VaultGemma 1B, a 1 billion parameter model within the Gemma family, fully trained with differential privacy. Pretrained on the identical data mixture used for the Gemm…

cs.LG2025

Urania: Differentially Private Insights into AI Use

Daogao Liu, Edith Cohen, Badih Ghazi +8

We introduce , a novel framework for generating insights about LLM chatbot interactions with rigorous differential privacy (DP) guarantees. The framework employs a private…

cs.CL2025

FlexOlmo: Open Language Models for Flexible Data Use

Weijia Shi, Akshita Bhagia, Kevin Farhat +20

We introduce FlexOlmo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained…

cs.LG2025

Linear-Time User-Level DP-SCO via Robust Statistics

Badih Ghazi, Ravi Kumar, Daogao Liu +1

User-level differentially private stochastic convex optimization (DP-SCO) has garnered significant attention due to the paramount importance of safeguarding user privacy in modern…

cs.CR2024

Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model Accuracy

Yangsibo Huang, Daogao Liu, Lynn Chua +7

Machine unlearning algorithms, designed for selective removal of training data from models, have emerged as a promising approach to growing privacy concerns. In this work, we expos…