Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
Csaba Dékány, Stefan Balauca, Robin Staab +2
Despite recent efforts in Large Language Model (LLM) safety and alignment, current adversarial attacks on frontier LLMs can still consistently force harmful generations. Although a…
cs.LG2024
CTBENCH: A Library and Benchmark for Certified Training
Yuhao Mao, Stefan Balauca, Martin Vechev
Training certifiably robust neural networks is an important but challenging task. While many algorithms for (deterministic) certified training have been proposed, they are often ev…
cs.LG2024
Gaussian Loss Smoothing Enables Certified Training with Tight Convex Relaxations
Stefan Balauca, Mark Niklas Müller, Yuhao Mao +3
Training neural networks with high certified accuracy against adversarial examples remains an open challenge despite significant efforts. While certification methods can effectivel…