Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Colin White, Samuel Dooley, Manley Roberts +15
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render ben…
cs.CL2025
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
Mayuka Jayawardhana, Renbo, Samuel Dooley +6
Large language models (LLMs) perform remarkably well on tabular datasets in zero- and few-shot settings, since they can extract meaning from natural language column headers that de…