120 citations · 196 across the 9 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
Cristina Garbacea, Heran Wang, Chenhao Tan
With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important chal…
cs.AI2025★ 1 cited
Evaluating the Goal-Directedness of Large Language Models
Tom Everitt, Cristina Garbacea, Alexis Bellot +4
To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require in…