paper

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

arXiv:2606.29091

Abstract

Tabular foundation models cannot reason about data produced by running systems without access to the rules that govern them. We make this statement falsifiable. The \emph{Operational Turing Test} (OTT) constructs pairs of legal and rule-violating database states whose - and -way column-value marginals match to a total variation of ; Le~Cam's lemma then bounds any values-only classifier at Bayes error. Three values-only baselines (XGBoost, TabICL, TabPFN) hit the bound exactly (accuracy , pre-registered two one-sided tests (TOST) ), raw row-level access does not help, exposing relational value consistency closes most of the gap, and only a classifier fed by seven executable rule-derived audits reaches classification accuracy. In three matched -state frontier large-language-model (LLM) runs, models given the schema, trigger source, rule tables, and state files classify at most legal states as LEGAL; GPT-5.5 accepts legal states even with higher reasoning effort and a Structured Query Language (SQL) executor. The access-ladder pattern also appears on a second schema with structurally distinct rule families (banking ledger: cross-row balance, cumulative aggregate). The barrier is identifiability, not capacity: scale, data, and richer features cannot cross it without operational grounding.

Accepted at the 2nd ICML Workshop on Foundation Models for Structured Data, 2026

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models · wovepaper