activity
20242026
most citedExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

3 citations · 3 across the 3 of their papers we have counts for

collaborators
Showing cs.CRShow all

5 papers · 1 filter

cs.CR2026

Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations

Aniket Anand, Yiwei Hou, Daniel Fields +4

This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. We design and use this benchmar…

cs.CR20263 cited

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Zhun Wang, Nico Schiller, Hongwei Li +13

AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity, making rigorous evaluation urgent. A critical capability is exploitation: turning a vulne…

cs.CR2026

DROIDCCT: Cryptographic Compliance Test via Trillion-Scale Measurement

Daniel Moghimi, Alexandru-Cosmin Mihai, Borbala Benko +6

We develop DroidCCT, a distributed test framework to evaluate the scale of a wide range of failures/bugs in cryptography for end users. DroidCCT relies on passive analysis of artif…

cs.CR2025

Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks

Milad Nasr, Yanick Fratantonio, Luca Invernizzi +7

As deep learning models become widely deployed as components within larger production systems, their individual shortcomings can create system-level vulnerabilities with real-world…

cs.CR2024

Magika: AI-Powered Content-Type Detection

Yanick Fratantonio, Luca Invernizzi, Loua Farah +9

The task of content-type detection -- which entails identifying the data encoded in an arbitrary byte sequence -- is critical for operating systems, development, reverse engineerin…