activity
20202026
most citedImproved Segmentation and Detection Sensitivity of Diffusion-Weighted Brain Infarct Lesions with Synthetically Enhanced Deep Learning

28 citations · 73 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL20262 cited

Reasoning Models Generate Societies of Thought

Junsol Kim, Shiyang Lai, Nino Scherrer +2

Large language models have achieved remarkable capabilities across domains, yet mechanisms underlying sophisticated reasoning remain elusive. Recent reasoning models outperform com…

cs.CL2025

Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis

Ferdinand Kapl, Emmanouil Angelis, Tobias Höppe +4

Gradually growing the depth of Transformers during training can not only reduce training cost but also lead to improved reasoning performance, as shown by MIDAS (Saunshi et al., 20…

cs.CL2025

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

Maty Bohacek, Nino Scherrer, Nicholas Dufour +3

The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas wher…

cs.CL20251 cited

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

Flor Miriam Plaza-del-Arco, Paul Röttger, Nino Scherrer +3

Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent wor…

cs.CL20245 cited

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97

This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…

cs.CL202311 cited

FinanceBench: A New Benchmark for Financial Question Answering

Pranab Islam, Anand Kannappan, Douwe Kiela +3

FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). It comprises 10,231 questions about publicly t…