3 papers
cs.AI2026
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
Elaine Lau, Markus Dücker, Ronak Chaudhary +24
Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…
cs.CL2026
Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction
Hui Wen Goh, Jonas Mueller
Structured Outputs from current LLMs exhibit sporadic errors, hindering enterprise AI deployment. We present CONSTRUCT, a real-time uncertainty estimator that scores the trustworth…
cs.LG2023
ActiveLab: Active Learning with Re-Labeling by Multiple Annotators
Hui Wen Goh, Jonas Mueller
In real-world data labeling applications, annotators often provide imperfect labels. It is thus common to employ multiple annotators to label data with some overlap between their e…