5 papers
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
Elaine Lau, Markus Dücker, Ronak Chaudhary +24
Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…
An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
Andreas Plesner, Francisco Guzmán, Anish Athalye
Reinforcement Learning with Verifiable Rewards (RLVR) has become a prominent method for post-training Large Language Models (LLMs). However, verifiers are rarely error-free; even d…
Translation as a Scalable Proxy for Multilingual Evaluation
Sheriff Issaka, Erick Rosas Gonzalez, Lieqi Liu +6
The rapid proliferation of LLMs has created a critical evaluation paradox: while LLMs claim multilingual proficiency, comprehensive non-machine-translated benchmarks exist for fewe…
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
H-, He-like recombination spectra -- V: On the dependence of the simulated line intensities on the number of electronic levels of the atoms
F. Guzmán, M. Chatzikos, G. Ferland
This paper presents a study of the dependence of the simulated intensities of recombination lines from hydrogen and helium atoms on the number of -resolved principal quantum…