2 papers
cs.AI2026
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
Matthew Kowal, Goncalo Paulo, Louis Jaburi +6
As large language models are increasingly trained and fine-tuned, practitioners need methods to identify which training data drive specific behaviors, particularly unintended ones.…
cs.CL2025
STACK: Adversarial Attacks on LLM Safeguard Pipelines
Ian R. McKenzie, Oskar J. Hollinsworth, Tom Tseng +5
Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic and OpenAI guard their latest Opus 4 model and GPT-5 mode…