Publications (19)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
Christine Ye, Sihan Yuan, Suchetha Cooray +10
Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use age…
Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration
LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz +63
The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that cha…
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
COMPOSITE-Stem
Kyle Waters, Lucas Nuzzi, Tadhg Looram +20
AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have prove…
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…
OpenThoughts-Agent: Data Recipes for Agentic Models
Negin Raoof, Richard Zhuang, Marianna Nezhurina +47
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts…
Investigating the Dark Energy Constraint from Strongly Lensed AGN at LSST-Scale
Sydney Erickson, Martin Millon, Padmavathi Venkatraman +12
The paper presents a scalable hierarchical inference framework to jointly analyze hundreds of strongly lensed AGN time delays from LSST, forecasting a ~2.5% measurement of H0 and a…
SLSim: a strong lensing population simulation package
Narayan Khadka, Simon Birrer, Henry Best +45
Gravitational lensing offers unique insights into cosmology by bending light around massive objects. Strong gravitational lensing, in particular, produces magnified and often multi…
Building Machine Learning Challenges for Anomaly Detection in Science
Elizabeth G. Campolongo, Yuan-Tang Chou, Ekaterina Govorkova +148
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals
Tadhg Looram, Lucas Nuzzi, Kyle Waters +1
Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward signals for post-training. Yet,…
Hyperluminous Supersoft X-Ray Sources in the Chandra Catalog
Andrea Sacchi, Kevin Paggeot, Steven Dillmann +2
Hyperluminous supersoft X-ray sources, such as bright extragalactic sources characterized by particularly soft X-ray spectra, offer a unique opportunity to study accretion onto sup…
The Advanced X-ray Imaging Satellite (AXIS) Community Science Book
Michael Koss, Nafisa Aftab, Steven W. Allen +395
The AXIS Community Science Book represents the collective effort of 592 scientists worldwide to define the transformative science enabled by the Advanced X-ray Imaging Satellite (A…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
A Poisson Process AutoDecoder for X-ray Sources
Yanke Song, Victoria Ashley Villar, Juan Rafael Martinez-Galarza +1
X-ray observing facilities, such as the Chandra X-ray Observatory and the eROSITA, have detected millions of astronomical sources associated with high-energy phenomena. The arrival…
Lens Model Accuracy in the Expected LSST Lensed AGN Sample
Padmavathi Venkatraman, Sydney Erickson, Phil Marshall +13
Strong gravitational lensing of active galactic nuclei (AGN) enables measurements of cosmological parameters through time-delay cosmography (TDC). With data from the upcoming LSST…
Learning Representations of Event Time Series with Sparse Autoencoders for Anomaly Detection, Similarity Search, and Unsupervised Classification
Steven Dillmann, Juan Rafael MartÃnez-Galarza
Event time series are sequences of discrete events occurring at irregular time intervals, each associated with a domain-specific observational modality. They are common in domains…
Representation Learning for Time-Domain High-Energy Astrophysics: Discovery of Extragalactic Fast X-ray Transient XRT 200515
Steven Dillmann, Juan Rafael MartÃnez-Galarza, Roberto Soria +2
We present a novel representation learning method for downstream tasks like anomaly detection, unsupervised classification, and similarity searches in high-energy data sets. This e…
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen +75
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…