2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Eungyeup Kim, Chenchen Gu, Vashisth Tiwari +1
While existing benchmarks demonstrate the near-perfect performance of large language models (LLMs) on various tasks, this apparent saturation often obscures the need for rigorous e…