1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2025
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
Haoyu Dong, Pengkun Zhang, Yan Gao +6
We introduce FinWorkBench (a.k.a. Finch) for evaluating AI agents on real-world, enterprise-grade finance and accounting workflows that interleave data entry, structuring, formatti…
cs.RO2025★ 1 cited
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
Adina Yakefu, Bin Xie, Chongyang Xu +34
Testing on real machines is indispensable for robotic control algorithms. In the context of learning-based algorithms, especially VLA models, demand for large-scale evaluation, i.e…