◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Geoffrey Irving

11 papers hereh-index 8328 citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4
  • last author6

Across the 10 of 11 papers where every author was matched, so the position is known.

fields
  • cs.AI5
  • cs.CY3
  • cs.LG2
  • cs.CR1
same name
  • Geoffrey Irving — 4 papers, h 2
  • Geoffrey Irving — 3 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedChain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

2 citations · 4 across the 9 of their papers we have counts for

collaborators
Showing cs.AIShow all

4 papers · 1 filter

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.AI2025

An alignment safety case sketch based on debate

Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton +1

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feed…

cs.AI2025

How to evaluate control measures for LLM agents? A trajectory from today to superintelligence

Tomek Korbak, Mikita Balesni, Buck Shlegeris +1

As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…

cs.AI2025

A sketch of an AI control safety case

Tomek Korbak, Joshua Clymer, Benjamin Hilton +2

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.