benchmarking 1data persistence 1evaluation metrics 1large language models 1LLM agents 1procedural agents 1program synthesis 1runtime systems 1safety-critical AI 1storage footprint 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents
Chenglin Yu, Li Yin, Ying Yu +4
The paper proposes compiling standard operating procedures into executable pseudo‑code and running them with a program‑guided stack machine that pages the active frame while a larg…
cs.AI2026
The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
Chenglin Yu, Hongquan Gui, Ying Yu +3
The paper introduces AgentFootprint, a benchmark that measures the persistent storage footprint left by large language model agents after execution, providing metrics on retention,…