most citedCybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

1 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CR2026

What Does It Mean to Break a Distillation Defense?

Lena Libon, Pura Peetathawatchai, Michael Aerni +2

Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work p…

cs.CR2026

Laundering AI Authority with Adversarial Examples

Jie Zhang, Pura Peetathawatchai, Florian Tramèr +1

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly…

cs.CV2024

Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings

Pura Peetathawatchai, Wei-Ning Chen, Berivan Isik +2

Personalizing large-scale diffusion models poses serious privacy risks, especially when adapting to small, sensitive datasets. A common approach is to fine-tune the model using dif…

cs.CR2024★ 1 cited

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Andy K. Zhang, Neil Perry, Riya Dulepet +24

Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policyma…

cs.LG2024★ 1 cited

On Fairness of Low-Rank Adaptation of Large Models

Zhoujie Ding, Ken Ziyu Liu, Pura Peetathawatchai +2

Low-rank adaptation of large models, particularly LoRA, has gained traction due to its computational efficiency. This efficiency, contrasted with the prohibitive costs of full-mode…