3 papers
cs.LG2025
Predicting the Performance of Black-box LLMs through Follow-up Queries
Dylan Sam, Marc Finzi, J. Zico Kolter
Reliably predicting the behavior of language models -- such as whether their outputs are correct or have been adversarially manipulated -- is a fundamentally challenging task. This…
cs.AI2025
Antidistillation Sampling
Yash Savani, Asher Trockman, Zhili Feng +5
Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability, model own…
cs.CV2024
Diffusing Differentiable Representations
Yash Savani, Marc Finzi, J. Zico Kolter
We introduce a novel, training-free method for sampling differentiable representations (diffreps) using pretrained diffusion models. Rather than merely mode-seeking, our method ach…