Sparse PPMI Graph Averaging for Random Indexing Embeddings
arXiv:2608.05724
Abstract
We study a specific sparse post-processing pipeline for Random Indexing (RI) on kinship analogies in a small fairytales corpus. The published artifacts use uniform RI context accumulation with 200 dimensions and eight nonzeros, followed by one residual graph average, , where is a row-normalized PPMI graph and . Terminal row normalization and per-dimension median/IQR scaling are then applied. On the Google analogy benchmark's family section, 272 of 506 questions are valid for every seed. Across five paired seeds, the complete pipeline raises accuracy from 19.41\% to 30.74\%, a gain of 11.32 percentage points with a nested-bootstrap 95\% confidence interval of [6.93, 15.89]. Robust scaling alone contributes 3.24 points [1.25, 5.38], while graph averaging without robust scaling contributes 6.18 points [2.63, 9.92]. A separate 40-question general grid does not support a general improvement: the full pipeline changes accuracy by -6.00 points [-13.50, -0.50], and averaging without robust scaling changes it by -6.50 points [-14.50, -0.50]. The supported positive claim is therefore limited to the covered fairytales kinship analogy set; the results do not establish a generally effective embedding method.