activity
20242026
collaborators

35 papers

cs.LG2026

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

Renfei Zhang, Niloofar Mireshghallah

Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here w…

cs.CR2026

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing

Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram

This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models. LeakyLMs is the first to demons…

cs.AI2026

Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics

Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah +3

What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data, objectives, architectures,…

cs.AI2026

SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

Kevin Han, Renfei Zhang, Kathy Wei +3

LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across div…

cs.LG2026

Boundary-targeted Membership Inference Attacks on Safety Classifiers

Anthony Hughes, Alexander Goldberg, Prince Jha +3

Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with large language models. Despit…

cs.CY2026

Muse Spark Safety & Preparedness Report

Cristina Menghini, Peter Ney, Hamza Kwisaba +117

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…