3 citations · 3 across the 3 of their papers we have counts for
3 papers
Office Comprehension Benchmark
Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.do…
InstructExcel: A Benchmark for Natural Language Instruction in Excel
Justin Payan, Swaroop Mishra, Mukul Singh +7
With the evolution of Large Language Models (LLMs) we can solve increasingly more complex NLP tasks across various domains, including spreadsheets. This work investigates whether L…
Co-audit: tools to help humans double-check AI-generated content
Andrew D. Gordon, Carina Negreanu, José Cambronero +9
Users are increasingly being warned to check AI-generated content for correctness. Still, as LLMs (and other generative models) generate more complex output, such as summaries, tab…