papers

Publications (19)

cs.CL2025

ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?

Christine Ye, Sihan Yuan, Suchetha Cooray +10

Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use age…

astro-ph.IM2026

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz +63

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that cha…

cs.SE2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…

cs.AI2026

COMPOSITE-Stem

Kyle Waters, Lucas Nuzzi, Tadhg Looram +20

AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have prove…

cs.SE2026

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Rishi Desai, Jesse Hu, Joan Cabezas +23

AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…

cs.AI2026

OpenThoughts-Agent: Data Recipes for Agentic Models

Negin Raoof, Richard Zhuang, Marianna Nezhurina +47

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts…

astro-ph.CO2026

Investigating the Dark Energy Constraint from Strongly Lensed AGN at LSST-Scale

Sydney Erickson, Martin Millon, Padmavathi Venkatraman +12

The paper presents a scalable hierarchical inference framework to jointly analyze hundreds of strongly lensed AGN time delays from LSST, forecasting a ~2.5% measurement of H0 and a…

#strong lensing#time-delay cosmography#dark energy#LSST
astro-ph.CO2026

SLSim: a strong lensing population simulation package

Narayan Khadka, Simon Birrer, Henry Best +45

Gravitational lensing offers unique insights into cosmology by bending light around massive objects. Strong gravitational lensing, in particular, produces magnified and often multi…

cs.LG2025

Building Machine Learning Challenges for Anomaly Detection in Science

Elizabeth G. Campolongo, Yuan-Tang Chou, Ekaterina Govorkova +148

Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not…

cs.AI2026

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…

cs.LG2026

AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals

Tadhg Looram, Lucas Nuzzi, Kyle Waters +1

Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward signals for post-training. Yet,…

astro-ph.HE2025

Hyperluminous Supersoft X-Ray Sources in the Chandra Catalog

Andrea Sacchi, Kevin Paggeot, Steven Dillmann +2

Hyperluminous supersoft X-ray sources, such as bright extragalactic sources characterized by particularly soft X-ray spectra, offer a unique opportunity to study accretion onto sup…

astro-ph.HE2026

The Advanced X-ray Imaging Satellite (AXIS) Community Science Book

Michael Koss, Nafisa Aftab, Steven W. Allen +395

The AXIS Community Science Book represents the collective effort of 592 scientists worldwide to define the transformative science enabled by the Advanced X-ray Imaging Satellite (A…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

astro-ph.IM2025

A Poisson Process AutoDecoder for X-ray Sources

Yanke Song, Victoria Ashley Villar, Juan Rafael Martinez-Galarza +1

X-ray observing facilities, such as the Chandra X-ray Observatory and the eROSITA, have detected millions of astronomical sources associated with high-energy phenomena. The arrival…

astro-ph.CO2025

Lens Model Accuracy in the Expected LSST Lensed AGN Sample

Padmavathi Venkatraman, Sydney Erickson, Phil Marshall +13

Strong gravitational lensing of active galactic nuclei (AGN) enables measurements of cosmological parameters through time-delay cosmography (TDC). With data from the upcoming LSST…

cs.LG2025

Learning Representations of Event Time Series with Sparse Autoencoders for Anomaly Detection, Similarity Search, and Unsupervised Classification

Steven Dillmann, Juan Rafael Martínez-Galarza

Event time series are sequences of discrete events occurring at irregular time intervals, each associated with a domain-specific observational modality. They are common in domains…

astro-ph.HE2025

Representation Learning for Time-Domain High-Energy Astrophysics: Discovery of Extragalactic Fast X-ray Transient XRT 200515

Steven Dillmann, Juan Rafael Martínez-Galarza, Roberto Soria +2

We present a novel representation learning method for downstream tasks like anomaly detection, unsupervised classification, and similarity searches in high-energy data sets. This e…

cs.AI2026

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Xiangyi Li, Yimin Liu, Wenbo Chen +75

Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…