Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
arXiv:2601.15322
Abstract
Tool-using agents can repeat a final decision while changing their recorded execution. We introduce the Determinism-Faithfulness Assurance Harness (DFAH), a framework that distinguishes decision repeatability, trajectory agreement, and evidence-conditioned faithfulness. Task correctness requires separately qualified labels and evaluation; evidence-conditioned faithfulness was not evaluated in the historical v2 agentic experiments. The original v2 study reported 4,705 agentic runs in three synthetic financial tasks and a decision-determinism/task-label-match correlation of r = -0.11 across 21 model-benchmark configuration summaries. This statistic is reproducible from the historical configuration table, but includes a subsequently excluded portfolio fixture. It is retained as a historical description, not evidence of statistical independence, predictive uselessness, or an architectural determinism-accuracy tradeoff. Recorded decision concentration and tool-path variation do not identify hidden model strategy. This correction qualifies the historical evidence and removes the deployment recommendations derived from those unsupported interpretations. A separate corrected study, DFAH-Bench (arXiv:2607.20491), provides qualified evidence of decision/path disagreement. The contribution retained here is a measurement framework: repeatability, observable execution, evidence alignment, and correctness require distinct evidence, with explicit capture and study boundaries.
23 pages, 4 figures, 8 tables. Corrects interpretation of historical results and clarifies study boundaries. Separate DFAH-Bench manuscript: arXiv:2607.20491. Code and data: https://github.com/ibm-client-engineering/output-drift-financial-llms. Original version accepted at the 2nd ICLR Workshop on Advances in Financial AI: Towards Agentic and Responsible Systems (ICLR 2026)