4 papers
Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval
Jeffrey Flynt
Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief store can achieve perfect recall and mask…
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Jeffrey Flynt
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fe…
OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
Jeffrey Flynt
Building and evaluating enterprise AI systems requires synthetic organizational corpora that are internally consistent, temporally structured, and cross-artifact traceable. Existin…
OrgForge-IT: A Verifiable Synthetic Benchmark for LLM-Based Insider Threat Detection
Jeffrey Flynt
Synthetic insider threat benchmarks face a consistency problem: corpora generated without an external factual constraint cannot rule out cross-artifact contradictions. The CERT dat…