papers

Publications (9)

cs.LG2026

Synthetic Persona Pretraining: Alignment from Token Zero

Julian Minder, Viktor Moskvoretskii, Raghav Singhal +12

As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant…

cs.CL2025

Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs

Jan Betley, Jorio Cocola, Dylan Feng +4

LLMs are useful because they generalize so well. But can you have too much of a good thing? We show that a small amount of finetuning in narrow contexts can dramatically shift beha…

cs.GT2022

A Framework for Single-Item NFT Auction Mechanism Design

Jason Milionis, Dean Hirsch, Andy Arditi +1

Lately, Non-Fungible Tokens (NFTs), i.e., uniquely discernible assets on a blockchain, have skyrocketed in popularity by addressing a broad audience. However, the typical NFT aucti…

cs.CL2026

Where Do Reasoning Models Refuse?

Kureha Yamaguchi, Benjamin Etheridge, Andy Arditi

Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response token. Reasoning models, by contrast, pr…

cs.CL2026

Real-Time Detection of Hallucinated Entities in Long-Form Generation

Oscar Obeso, Andy Arditi, Javier Ferrando +3

Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing halluc…

cs.CL2025

Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Runjin Chen, Andy Arditi, Henry Sleight +2

Large language models interact with users through a simulated 'Assistant' persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviat…