3 papers
cs.SE2026
The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks
Bardia Mohammadi, Lars Klein, Aman Chadha +2
Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a c…
cs.CL2026
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Fay Elhassan, David Sasu, Alexandra Kulinkina +2
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive O…
cs.CL2025
TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
David Schulmeister, Valentin Hartmann, Lars Klein +1
Today, a lot of research on language models is focused on large, general-purpose models. However, many NLP pipelines only require models with a well-defined, small set of capabilit…