From the 1 of 2 linked papers with an AI index.
2 papers
cs.SE2026
Specula: Scaling formal specifications for autonomous model checking of system code
Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang +6
Specula is an autonomous system that uses large language model agents to generate TLA+ specifications for complex system code and then applies model checking to discover bugs.
cs.AI2026
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial +5
AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to…