3 papers
cs.CL2026
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Congmin Zheng, Jiachen Zhu, Jianghao Lin +6
Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…
cs.AI2025
Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
Mathieu Andreux, Märt Bakler, Yanael Barbier +50
Building agents that generalize across web, desktop, and mobile environments remains an open challenge, as prior systems rely on environment-specific interfaces that limit cross-pl…
cs.AI2025
Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights
Mathieu Andreux, Breno Baldas Skuk, Hamza Benchekroun +41
We present Surfer-H, a cost-efficient web agent that integrates Vision-Language Models (VLM) to perform user-defined tasks on the web. We pair it with Holo1, a new open-weight coll…