Grade Encoding and the Structural Representation of Student Academic Trajectories
arXiv:2606.12946
Abstract
Does the conversion of academic assessment from percentage scores to letter grades merely represent an adjustment in information precision, or does it systematically alter the underlying structure of student academic data? Drawing on 68 mathematics exam scores from 75 primary school students, this study employs Encoding Transformation simulation to compare structural differences in the same dataset under two encoding schemes across three dimensions: information loss, distance structure change, and clustering stability. Results indicate that letter-grade encoding compresses the mean pairwise distance in the trajectory feature space from 20.50 to 1.06 (a compression ratio of approximately 19:1); after standardization, the density gradient of the distance distribution is systematically flattened, with kurtosis decreasing by 0.54 and the coefficient of variation decreasing by 0.16; and the clustering structure becomes highly sensitive to minor sample perturbations-removing a single extreme observation causes the optimal K to jump from 4 to 8 and clustering reproducibility to plummet from 95% to 62%. Concurrently, a Compression Paradox is observed: letter-grade encoding increases the Silhouette coefficient while simultaneously reducing clustering stability. These findings demonstrate that letter-grade encoding systematically alters the structural representation of longitudinal academic data and exerts measurable effects on clustering stability.
3 figures, 12 tables