◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Xander Davies

17 papers hereh-index 101.5k citations26 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author11
  • last author4

Across the 17 of 17 papers where every author was matched, so the position is known.

fields
  • cs.LG8
  • cs.AI3
  • cs.CR3
  • cs.CL2
  • cs.SE1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.AIShow all

3 papers · 1 filter

cs.AI2026

Evaluating whether AI models would sabotage AI safety research

Robert Kirk, Alexandra Souly, Kai Fronsdal +2

We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two co…

cs.AI2026

UK AISI Alignment Evaluation Case-Study

Alexandra Souly, Robert Kirk, Jacob Merizian +2

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate…

cs.AI2025

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Andy Zou, Maxwell Lin, Eliot Jones +14

Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these sys…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.