28 citations · 73 across the 11 of their papers we have counts for
8 papers · 1 filter
Reasoning Models Generate Societies of Thought
Junsol Kim, Shiyang Lai, Nino Scherrer +2
Large language models have achieved remarkable capabilities across domains, yet mechanisms underlying sophisticated reasoning remain elusive. Recent reasoning models outperform com…
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
Ferdinand Kapl, Emmanouil Angelis, Tobias Höppe +4
Gradually growing the depth of Transformers during training can not only reduce training cost but also lead to improved reasoning performance, as shown by MIDAS (Saunshi et al., 20…
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
Maty Bohacek, Nino Scherrer, Nicholas Dufour +3
The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas wher…
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
Flor Miriam Plaza-del-Arco, Paul Röttger, Nino Scherrer +3
Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent wor…
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97
This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…
FinanceBench: A New Benchmark for Financial Question Answering
Pranab Islam, Anand Kannappan, Douwe Kiela +3
FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). It comprises 10,231 questions about publicly t…