5 papers · 1 filter
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
Nicole Mitchell, Dhruv Agarwal, Maty Bohacek +2
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features…
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
The DeepSpeak-Agentic Dataset
Sarah Barrington, Maty Bohacek, Hany Farid
We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We use this dataset to evaluat…
Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
Matyas Bohacek, Ignacio Vilanova Echavarri
Generative Artificial Intelligence (GAI) has experienced exponential growth in recent years, partly facilitated by the abundance of large-scale open-source datasets. These datasets…