2 papers
cs.CL2025
Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
Sabrina Sadiekh, Elena Ericheva, Chirag Agarwal
Advances in unsupervised probes such as Contrast-Consistent Search (CCS), which reveal latent beliefs without relying on token outputs, raise the question of whether these methods…
cs.LG2025
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Hjalmar Wijk, Tao Lin, Joel Becker +20
Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations fo…