6 papers
False Prophets: On the Security of World Models in Agentic Systems
Erik Imgrund, Anna Wimbauer, Klim Kireev +1
Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent…
When a Zero-Shooter Cheats: Improving Age Estimation via Activation Steering
Erik Imgrund, Pia Hanfeld, Klim Kireev +1
Different age-related regulations have been proposed to protect minors from harmful content and interactions online. Automated age estimation is central to enforcing such regulatio…
Toward Securing AI Agents Like Operating Systems
Lukas Pirch, Micha Horlboge, Patrick GroÃmann +4
Autonomous agents based on large language models (LLMs) are rapidly emerging as a general-purpose technology, with recent systems such as OpenClaw extending their capabilities thro…
Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models
Ana-Maria Cretu, Klim Kireev, Amro Abdalla +5
We evaluate the effectiveness of filtering child images from training datasets of text-to-image models to prevent model misuse to create child sexual abuse material (CSAM). First,…
A Manually Annotated Image-Caption Dataset for Detecting Children in the Wild
Klim Kireev, Ana-Maria Creţu, Raphael Meier +3
Platforms and the law regulate digital content depicting minors (defined as individuals under 18 years of age) differently from other types of content. Given the sheer amount of co…
Inverting Black-Box Face Recognition Systems via Zero-Order Optimization in Eigenface Space
Anton Razzhigaev, Matvey Mikhalchuk, Klim Kireev +3
Reconstructing facial images from black-box recognition models poses a significant privacy threat. While many methods require access to embeddings, we address the more challenging…