Publications (9)
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder, Viktor Moskvoretskii, Raghav Singhal +12
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant…
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
Jan Betley, Jorio Cocola, Dylan Feng +4
LLMs are useful because they generalize so well. But can you have too much of a good thing? We show that a small amount of finetuning in narrow contexts can dramatically shift beha…
A Framework for Single-Item NFT Auction Mechanism Design
Jason Milionis, Dean Hirsch, Andy Arditi +1
Lately, Non-Fungible Tokens (NFTs), i.e., uniquely discernible assets on a blockchain, have skyrocketed in popularity by addressing a broad audience. However, the typical NFT aucti…
Where Do Reasoning Models Refuse?
Kureha Yamaguchi, Benjamin Etheridge, Andy Arditi
Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response token. Reasoning models, by contrast, pr…
Real-Time Detection of Hallucinated Entities in Long-Form Generation
Oscar Obeso, Andy Arditi, Javier Ferrando +3
Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing halluc…
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Runjin Chen, Andy Arditi, Henry Sleight +2
Large language models interact with users through a simulated 'Assistant' persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviat…