collaborators

5 papers

cs.LG2026

Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness

Tavish McDonald, Bo Lei, Stanislav Fort +2

Test-time reasoning has raised benchmark performances and even shown promise in addressing the historically intractable problem of making models robust to adversarially out-of-dist…

cs.LG2026

Solving adversarial examples requires solving exponential misalignment

Alessandro Salvatore, Stanislav Fort, Surya Ganguli

Adversarial attacks - input perturbations imperceptible to humans that fool neural networks - remain both a persistent failure mode in machine learning, and a phenomenon with myste…

cs.CV2026

Representations of Text and Images Align From Layer One

Evžen Wybitul, Javier Rando, Florian Tramèr +1

We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…

cs.CV2025

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models

Stanislav Fort, Jonathan Whitaker

We demonstrate that discriminative models inherently contain powerful generative capabilities, challenging the fundamental distinction between discriminative and generative archite…

cs.CR2025

A Note on Implementation Errors in Recent Adaptive Attacks Against Multi-Resolution Self-Ensembles

Stanislav Fort

This note documents an implementation issue in recent adaptive attacks (Zhang et al. [2024]) against the multi-resolution self-ensemble defense (Fort and Lakshminarayanan [2024]).…