paper

RSD: Moving Local Triangular Charts for Auditing Language-Model Hidden States

arXiv:2605.17482

Abstract

We study Relational Semantic Decomposition, abbreviated as RSD, as a moving local triangular chart audit for language-model hidden states. For repeated occurrences of one target word, RSD fits a shared three-anchor membership chart at layer or token-time . The hidden-state channel uses ; the invariant readout is the induced occurrence co-membership relation, and records what the fitted root chart leaves outside the chart. The broader joint audit reuses the same membership chart for relation data, , such as an attention-derived occurrence relation. The current GPT-2 evidence is the -channel hidden-state audit with Word-in-Context labels used as an external same-sense versus different-sense reference relation. On full WiC train, the root chart passes 16 of 53 eligible target words; this is audit coverage, not GPT-2 task accuracy. Token-time and pair-level diagnostics show the main regimes: \texttt{make} and \texttt{break} align at the target state, \texttt{drive} and \texttt{stay} improve after right context in small-count exploratory cases, and \texttt{play} remains a localized root-chart failure whose final same-sense pairs are not closer and have larger residual discrepancy. The resulting claim is diagnostic: RSD reports where a sense relation is visible in root co-membership and which failures become residual branch candidates or attention-channel obligations.

8 pages, 1 figure. Revised version with clarified scope, experiments, and limitations

RSD: Moving Local Triangular Charts for Auditing Language-Model Hidden States · wovepaper