3 papers
cs.LG2025
Training LLMs for Honesty via Confessions
Manas Joglekar, Jeremy Chen, Gabriel Wu +4
Large language models (LLMs) can be dishonest when reporting on their actions and beliefs -- for example, they may overstate their confidence in factual claims or cover up evidence…
cs.IT2025
Testing Tensor Products of Algebraic Codes
Sumegha Garg, Madhu Sudan, Gabriel Wu
Motivated by recent advances in locally testable codes and quantum LDPCs based on robust testability of tensor product codes, we explore the local testability of tensor products of…
cs.LG2025
Estimating the Probabilities of Rare Outputs in Language Models
Gabriel Wu, Jacob Hilton
We consider the problem of low probability estimation: given a machine learning model and a formally-specified input distribution, how can we estimate the probability of a binary p…