3 papers
cs.AI2026
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
Prajjwal Gupta, Prasang Gupta, Vishal Bhutani +4
As agentic LLM systems move from prototypes to deployment across increasingly diverse domains, evaluating them has become both more important and more difficult. The challenge is n…
cs.CL2026
Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
Anmol Gulati, Sahil Sen, Waqar Sarguroh +1
Recent advances in multimodal Retrieval-Augmented Generation (RAG) enable Large Language Models (LLMs) to analyze enterprise spreadsheet workbooks containing millions of cells, cro…
cs.CL2026
From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding
Anmol Gulati, Sahil Sen, Waqar Sarguroh +1
Large Language Models (LLMs) struggle to reason over large-scale enterprise spreadsheets containing thousands of numeric rows, multiple linked sheets, and embedded visual content s…