2 citations · 2 across the 3 of their papers we have counts for
10 papers
Mixing Times of Glauber Dynamics on Masked Language Models
Suvadip Sana, Sami Wolf, Neer Mehta +4
Masked language models (MLMs) define local conditional distributions over tokens but do not, in general, correspond to any consistent joint distribution over sequences. This raises…
Exploring the Impact of Dataset Statistical Effect Size on Model Performance and Data Sample Size Sufficiency
Arya Hatamian, Lionel Levine, Haniyeh Ehsani Oskouie +1
Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determine the adequacy of a dataset prio…
On Safety Risks in Experience-Driven Self-Evolving Agents
Weixiang Zhao, Yichen Zhang, Yingshuo Wang +8
Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduc…
EigenBench: A Comparative Behavioral Measure of Value Alignment
Jonathn Chang, Leonhard Piff, Suvadip Sana +2
Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for compara…
Exploring Cross-model Neuronal Correlations in the Context of Predicting Model Performance and Generalizability
Haniyeh Ehsani Oskouie, Sajjad Ghiasvand, Lionel Levine +1
As Artificial Intelligence (AI) models are increasingly integrated into critical systems, the need for a robust framework to establish the trustworthiness of AI is increasingly par…
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Elliot Glazer, Ege Erdil, Tamay Besiroglu +21
We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most…