Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
Si Chen, Le Huy Khiem, Annalisa Szymanski +3
Open-ended question answering (QA) evaluates a model's ability to perform contextualized reasoning beyond factual recall. This challenge is especially acute in practice-based domai…
cs.CL2025
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen +4
Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluat…