1 paper
Alexander Gill, Abhilasha Ravichander, Ana MarasoviÄ
Large language models (LLMs) are increasingly used for data generation. However, creating evaluation benchmarks raises the bar for this emerging paradigm. Benchmarks must target sp…