Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
Maty Bohacek, Nino Scherrer, Nicholas Dufour +3
The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas wher…
cs.CL2024
Machine Psychology
Thilo Hagendorff, Ishita Dasgupta, Marcel Binz +5
Large language models (LLMs) show increasingly advanced emergent capabilities and are being incorporated across various societal domains. Understanding their behavior and reasoning…