5 citations · 10 across the 15 of their papers we have counts for
3 papers · 1 filter
FragBench: Cross-Session Attacks Hidden in Benign-Looking Fragments
Astha Mehta, Niruthiha Selvanayagam, Cedric Lam +10
An attacker can split a malicious goal into sub-prompts that each look benign on their own and only become harmful in combination. Existing LLM safety benchmarks evaluate prompts o…
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
David Williams-King, Linh Le, Adam Oberman +1
As LLMs develop increasingly advanced capabilities, there is an increased need to minimize the harm that could be caused to society by certain model outputs; hence, most LLMs have…
XDA: Accurate, Robust Disassembly with Transfer Learning
Kexin Pei, Jonas Guan, David Williams-King +2
Accurate and robust disassembly of stripped binaries is challenging. The root of the difficulty is that high-level structures, such as instruction and function boundaries, are abse…