19 papers
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
Nicole Mitchell, Dhruv Agarwal, Maty Bohacek +2
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features…
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
The DeepSpeak-Agentic Dataset
Sarah Barrington, Maty Bohacek, Hany Farid
We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We use this dataset to evaluat…
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
Maty Bohacek, Nino Scherrer, Nicholas Dufour +3
The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas wher…
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
Stefan Baack, Christo Buschek, Maty Bohacek
The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases and company blog posts, whe…