4 papers
PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming
Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson +13
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based test…
JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
Sandip Ghoshal, Anshul Mittal, Jyotika Singh +9
Large language model (LLM) agents augmented with external tools often struggle as number of tools grow large and become domain-specific. In such settings, ambiguous tool descriptio…
Routesplain: Towards Faithful and Intervenable Routing for Software-related Tasks
Adam Štorek, Vikas Upadhyay, Marianne Menglin Liu +6
LLMs now tackle a wide range of software-related tasks, yet we show that their performance varies markedly both across and within these tasks. Routing user queries to the appropria…
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal +7
Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such…