paper

Optimal Allocation of Embedding Dimensions under Finite-Sample Constraints

arXiv:2608.24592

Abstract

The embedding dimension of categorical predictors is usually selected through heuristic tuning, although it directly affects model complexity, approximation quality, and finite-sample generalization. This paper formulates embedding dimension selection as a constrained allocation problem. The main contribution is to show that embedding capacity can be allocated across heterogeneous categorical predictors according to an explicit approximation-estimation tradeoff. We characterize approximation error through the singular-value tail of the latent category representation, while estimation error increases with total embedding complexity. Under a fixed global embedding budget and a tractable approximation model, this leads to a closed-form allocation rule in which the dimension assigned to each predictor is proportional to the square root of its approximation value relative to its parameter cost. Simulation experiments support the proposed approximation-estimation interpretation and show that the allocation rule improves budget efficiency relative to standard uniform and cardinality-based heuristics, particularly when the budget is binding and predictor heterogeneity is substantial. A real-data healthcare application further shows improvements in predictive accuracy and probabilistic calibration. Overall, the results establish embedding-dimension allocation as a principled finite-sample optimization problem rather than a purely heuristic modeling choice.

Keywords: Machine Learning in OR, Constrained Optimization, Nonlinear Programming, Multivariate Statistics, Predictive Models

Optimal Allocation of Embedding Dimensions under Finite-Sample Constraints · wovepaper