1 paper · 1 filter
Maya Patel, Aditi Anand
Benchmarking modern large language models (LLMs) on complex and realistic tasks is critical to advancing their development. In this work, we evaluate the factual accuracy and citat…