◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Chloe Li

3 papers hereh-index 225 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.CR1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.AI2026

Model Spec Midtraining: Improving How Alignment Training Generalizes

Chloe Li, Nevan Wichers, Sara Price +2

Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- trai…

cs.AI2026

Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives

Chloe Li, Mary Phuong, Daniel Tan

As AI systems become more capable of complex agentic tasks, they also become more capable of pursuing undesirable objectives and causing harm. Previous work has attempted to catch…

cs.CR2025

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

Chloe Li, Mary Phuong, Noah Y. Siegel

Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbaggin…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.