3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.AI2026
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
Elaine Lau, Markus Dücker, Ronak Chaudhary +24
Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…
cs.CV2023★ 3 cited
ObjectLab: Automated Diagnosis of Mislabeled Images in Object Detection Data
Ulyana Tkachenko, Aditya Thyagarajan, Jonas Mueller
Despite powering sensitive systems like autonomous vehicles, object detection remains fairly brittle in part due to annotation errors that plague most real-world training datasets.…