5 papers
MortarBench: Evaluating Mortgage Loan Origination Agents
Matthew Toles, Yunan Lu, Manav Munjal +6
Loan origination is the process by which a lender creates a new loan, from application and underwriting through approval and funding. This process serves a critical role in evaluat…
PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures
Yunan Lu, Luigi Liu, Omar Yahia +2
Evaluating LLM-based agents remains challenging because identifying meaningful failure cases often requires substantial human effort to design realistic test scenarios. Prior works…
AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions
Minghao Chen, Xinyi Hu, Zhou Yu +1
Large Language Model (LLM) based agents have demonstrated proficiency in multi-step interactions with graphical user interfaces (GUIs). While most research focuses on improving sin…
FormGym: Doing Paperwork with Agents
Matthew Toles, Rattandeep Singh, Isaac Song +1
Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without access to OCR, typeset PDF text, or a DOM.…
Program Synthesis Dialog Agents for Interactive Decision-Making
Matthew Toles, Nikhil Balwani, Rattandeep Singh +2
Many real-world eligibility problems, ranging from medical diagnosis to tax planning, can be mapped to decision problems expressed in natural language, wherein a model must make a…