2 papers
cs.SE2026
AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents
Laïla Elkoussy, Julien Perez
Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical settings, the procedure itsel…
cs.SE2026
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
Laïla Elkoussy, Julien Perez
In this paper, we introduce SWE-QA, a text and code corpus aimed at benchmarking multi-hop code comprehension, addressing the gap between simplified evaluation tasks and the comple…