4 papers
The Future of Facts: Tracing the Factual Generation-Verification Gap
Tim R. Davidson, Anja Surina, Caglar Gulcehre
Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-g…
Reasoning-Driven Synthetic Data Generation and Evaluation
Tim R. Davidson, Benoit Seguin, Enrico Bacis +2
Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with h…
Evaluating Language Model Agency through Negotiations
Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski +4
We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings o…
The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation
Tim R. Davidson, Adam Fourney, Saleema Amershi +3
The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with differen…