1 paper · 1 filter
Devleena Das, Rajeev Patwari, Vikram Kumar Bukka +3
Evaluating LLMs across many model variants -- quantized, fine-tuned, or deployment-specific -- requires running large benchmarks repeatedly, a process that can take tens of hours p…