2 papers
cs.SE2026
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
Francesco Dente, Dario Satriani, Paolo Papotti
Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adhere…
cs.CL2025
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
Dario Satriani, Enzo Veltri, Donatello Santoro +1
Factuality in Large Language Models (LLMs) is a persistent challenge. Current benchmarks often assess short factual answers, overlooking the critical ability to generate structured…