3 papers
cs.CL2024
Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
Peter Mühlbacher, Nikos I. Bosse, Lawrence Phillips
We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research…
math.PR2019
Critical Parameters for Loop and Bernoulli Percolation
Peter Mühlbacher
We consider a class of random loop models (including the random interchange process) that are parametrised by a time parameter . Intuitively, larger means more randomn…
math.PR2018
Bounds on the norm of Wigner-type random matrices
László Erdős, Peter Mühlbacher
We consider a Wigner-type ensemble, i.e. large hermitian random matrices with centered independent entries and with a general matrix of variances $S_{xy}=\mathb…