collaborators

11 papers

cs.CY2026

Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems

Susanne Gaube, Markus Langer, Tim Miller +17

The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that may only be ameliorated by human…

cs.CY2026

The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research

Xiaoyan Bai, Alexander Baumgartner, Haojia Sun +2

Reproducibility crises across sciences highlight the limitations of the paper-centric review system in assessing the rigor and reproducibility of research. AI agents that autonomou…

econ.GN2026

Human-AI Collaboration in Radiology: The Case of Pulmonary Embolism

Paul Goldsmith-Pinkham, Chenhao Tan, Alexander K. Zentefis

We study how radiologists use AI to diagnose pulmonary embolism (PE), tracking over 100,000 scans interpreted by nearly 400 radiologists during the staggered rollout of a real-worl…

cs.AI2025

Know Thyself? On the Incapability and Implications of AI Self-Recognition

Xiaoyan Bai, Aryan Shrivastava, Ari Holtzman +1

Self-recognition is a crucial metacognitive capability for AI systems, relevant not only for psychological analysis but also for safety, particularly in evaluative scenarios. Motiv…

cs.LG2025

Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls

Xiaoyan Bai, Itamar Pres, Yuntian Deng +5

Language models are increasingly capable, yet still fail at a seemingly simple task of multi-digit multiplication. In this work, we study why, by reverse-engineering a model that s…

cs.CL2025

MoVa: Towards Generalizable Classification of Human Morals and Values

Ziyu Chen, Junfei Sun, Chenxi Li +6

Identifying human morals and values embedded in language is essential to empirical studies of communication. However, researchers often face substantial difficulty navigating the d…