2 papers
cs.CL2025
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
Guanghui Wang, Jinze Yu, Xing Zhang +5
Large Language Models (LLMs) are increasingly deployed for structured data generation, yet output consistency remains critical for production applications. We introduce a comprehen…
cs.AI2025
SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems
Qian Cheng, Ruize Tang, Emilie Ma +7
Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in g…