binary outcome classification 1code-owned forecasting 1delayed ground truth evaluation 1large language models 1model auditing 1non-monotonic scoring 1omission exploitation 1plan evaluation 1strategic route generation 1typed-state gating 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation
Aleh Manchuliantsau
The paper shows that LLM‑generated strategic plans can be artificially improved by deleting intermediate steps, which raises the evaluator's score without real progress, and propos…
cs.AI2026
From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth
Aleh Manchuliantsau
The paper introduces RouteCast, a framework for evaluating model‑generated strategic routes when ground truth is delayed or private, using code‑owned provisional forecasts and repo…