Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
The Future of Facts: Tracing the Factual Generation-Verification Gap
Tim R. Davidson, Anja Surina, Caglar Gulcehre
Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-g…
cs.CL2026★ 4 cited
Evaluating Language Model Agency through Negotiations
Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski +4
We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings o…
cs.CL2024
Self-Recognition in Language Models
Tim R. Davidson, Viacheslav Surkov, Veniamin Veselovsky +3
A rapidly growing number of applications rely on a small set of closed-source language models (LMs). This dependency might introduce novel security risks if LMs develop self-recogn…