4 papers
LLM Jaggedness Unlocks Scientific Creativity
Shray Mathur, J. Anibal Boscoboinik, Esther H. R. Tsai +1
As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing unevenly across tasks, domains, an…
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
Seungone Kim, Dongkeun Yoon, Kiril Gashteovski +55
With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientis…
EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines
Noah van der Vleuten, Anthony Flores, Shray Mathur +4
Evaluating large language models (LLMs) for instrument control requires methods that go beyond standard, stateless algorithmic benchmarks, since the behavior of physical systems ca…
VISION: A Modular AI Assistant for Natural Human-Instrument Interaction at Scientific User Facilities
Shray Mathur, Noah van der Vleuten, Kevin Yager +1
Scientific user facilities, such as synchrotron beamlines, are equipped with a wide array of hardware and software tools that require a codebase for human-computer-interaction. Thi…