collaborators

5 papers

cs.AI2026

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

Elaine Lau, Markus Dücker, Ronak Chaudhary +24

Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…

cs.LG2026

An Imperfect Verifier is Good Enough: Learning with Noisy Rewards

Andreas Plesner, Francisco Guzmán, Anish Athalye

Reinforcement Learning with Verifiable Rewards (RLVR) has become a prominent method for post-training Large Language Models (LLMs). However, verifiers are rarely error-free; even d…

cs.CL2026

Translation as a Scalable Proxy for Multilingual Evaluation

Sheriff Issaka, Erick Rosas Gonzalez, Lieqi Liu +6

The rapid proliferation of LLMs has created a critical evaluation paradox: while LLMs claim multilingual proficiency, comprehensive non-machine-translated benchmarks exist for fewe…

cs.SE2026

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

Redacted by arXiv

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…

astro-ph.GA2025

H-, He-like recombination spectra -- V: On the dependence of the simulated line intensities on the number of electronic levels of the atoms

F. Guzmán, M. Chatzikos, G. Ferland

This paper presents a study of the dependence of the simulated intensities of recombination lines from hydrogen and helium atoms on the number of -resolved principal quantum…