1 paper · 1 filter
Liyan Chen, Yael Tauman Kalai, Zoe Xi
As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focus…