1 citations · 1 across the 2 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025★ 1 cited
Training LLMs for Honesty via Confessions
Manas Joglekar, Jeremy Chen, Gabriel Wu +4
Large language models (LLMs) can be dishonest when reporting on their actions and beliefs -- for example, they may overstate their confidence in factual claims or cover up evidence…
cs.LG2024
Estimating the Probabilities of Rare Outputs in Language Models
Gabriel Wu, Jacob Hilton
We consider the problem of low probability estimation: given a machine learning model and a formally-specified input distribution, how can we estimate the probability of a binary p…
cs.LG2024
Towards Understanding Human Emotional Fluctuations with Sparse Check-In Data
Sagar Paresh Shah, Ga Wu, Sean W. Kortschot +1
Data sparsity is a key challenge limiting the power of AI tools across various domains. The problem is especially pronounced in domains that require active user input rather than m…