NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (19)

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

cs.CL2025

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +1340

cs.CR2026

Phantom Transfer: Data Poisoning can Survive Data-Level Defences

Andrew Draganov, Tolga H. Dur, Anandmayi Bhongade +1

cs.AI2023

Model evaluation for extreme risks

Toby Shevlane, Sebastian Farquhar, Ben Garfinkel +18

cs.LG2026

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

Oliver Makins, Orazio Angelini, Zohreh Shams +1

cs.LG2022

Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals

Rohin Shah, Vikrant Varma, Ramana Kumar +4

cs.AI2025

CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring

Benjamin Arnav, Pablo Bernabeu-Pérez, Nathan Helm-Burger +3

cs.LG2024

Evaluating Frontier Models for Dangerous Capabilities

Mary Phuong, Matthew Aitchison, Elliot Catt +24

cs.AI2026

Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives

Chloe Li, Mary Phuong, Daniel Tan

cs.LG2021

Towards Understanding Knowledge Distillation

Mary Phuong, Christoph H. Lampert

cs.CY2025

From Stability to Inconsistency: A Study of Moral Preferences in LLMs

Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1

cs.LG2026

Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization

Frank Xiao, Mary Phuong

cs.LG2026

Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents

Frank Xiao, Mary Phuong

cs.LG2025

Evaluating Frontier Models for Stealth and Situational Awareness

Mary Phuong, Roland S. Zimmermann, Ziyue Wang +6

cs.LG2022

Formal Algorithms for Transformers

Mary Phuong, Marcus Hutter

cs.CR2026

GDM AI Control Roadmap

Mary Phuong, Erik Jenner, Laurent Simon +4

The paper presents the GDM AI Control Roadmap, a framework for internal security against potentially misaligned AI agents, including threat modeling, capability‑based mitigation ti…

#ai safety#threat modeling#defense tiers#capability-based mitigation
cs.CR2025

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

Chloe Li, Mary Phuong, Noah Y. Siegel

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38