Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Fay Elhassan, David Sasu, Alexandra Kulinkina +2
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive O…
cs.CL2025
TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
David Schulmeister, Valentin Hartmann, Lars Klein +1
Today, a lot of research on language models is focused on large, general-purpose models. However, many NLP pipelines only require models with a well-defined, small set of capabilit…
cs.CL2025
Fleet of Agents: Coordinated Problem Solving with Large Language Models
Lars Klein, Nearchos Potamitis, Roland Aydin +3
While numerous frameworks have been developed to enhance the reasoning abilities of large language models (LLMs), there is a scarcity of methods that effectively balance the trade-…