3 papers
cs.CL2025
The SMeL Test: A simple benchmark for media literacy in language models
Gustaf Ahdritz, Anat Kleiman
The internet is rife with unattributed, deliberately misleading, or otherwise untrustworthy content. Though large language models (LLMs) are often tasked with autonomous web browsi…
cs.LG2024
Provable Uncertainty Decomposition via Higher-Order Calibration
Gustaf Ahdritz, Aravind Gollakota, Parikshit Gopalan +2
We give a principled method for decomposing the predictive uncertainty of a model into aleatoric and epistemic components with explicit semantics relating them to the real-world da…
cs.LG2024
Modeling Real-Time Interactive Conversations as Timed Diarized Transcripts
Garrett Tanzer, Gustaf Ahdritz, Luke Melas-Kyriazi
Chatbots built upon language models have exploded in popularity, but they have largely been limited to synchronous, turn-by-turn dialogues. In this paper we present a simple yet ge…