1 paper · 1 filter
Ishir Garg, Neel Kolhe, Xuandong Zhao +1
Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to evaluate these capabilities rem…