collaborators

11 papers

cs.CY2026

Muse Spark Safety & Preparedness Report

Cristina Menghini, Peter Ney, Hamza Kwisaba +117

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…

cs.CL2026

On Safety Risks in Experience-Driven Self-Evolving Agents

Weixiang Zhao, Yichen Zhang, Yingshuo Wang +8

Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduc…

cs.LG2025

Progress over Points: Reframing LM Benchmarks Around Scientific Objectives

Alwin Jin, Sean M. Hendryx, Vaskar Nath

Current benchmarks that test LLMs on static, already-solved problems (e.g., math word problems) effectively demonstrated basic capability acquisition. The natural progression has b…

cs.SE2025

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Xiang Deng, Jeff Da, Edwin Pan +19

We introduce SWE-Bench Pro, a substantially more challenging benchmark that builds upon the best practices of SWE-BENCH [25], but is explicitly designed to capture realistic, compl…

cs.AI2025

ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents

Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi +13

Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities,…

cs.LG2025

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Anisha Gunjal, Anthony Wang, Elaine Lau +4

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it…