2 papers
cs.SE2025
EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines
Noah van der Vleuten, Anthony Flores, Shray Mathur +4
Evaluating large language models (LLMs) for instrument control requires methods that go beyond standard, stateless algorithmic benchmarks, since the behavior of physical systems ca…
cs.SE2025
Dr. Boot: Bootstrapping Program Synthesis Language Models to Perform Repairing
Noah van der Vleuten
Language models for program synthesis are usually trained and evaluated on programming competition datasets (MBPP, APPS). However, these datasets are limited in size and quality, w…