statistics

An association measure for mixed-type variables

arXiv:2607.26508

summary

The paper introduces a label‑invariant, normalized measure of association for a real‑valued variable and a categorical variable, along with a fast O(n log n) estimator and associated inference tools.

Abstract

Quantifying the association between a real-valued variable and a categorical variable is a fundamental task in data analysis. Existing methods often rely on parametric assumptions or arbitrary integer encoding, which may lead to unstable results. We propose a label-invariant population measure of association, , specifically designed for the mixed real-valued-categorical setting. The proposed measure is normalized between 0 and 1; it equals 0 if and only if the variables are independent and 1 if and only if the categorical variable is a measurable function of the real-valued one. We also introduce a corresponding sample estimator, , computable in time. These measures are invariant to permutations of category labels and strictly monotone transformations of the real-valued variable. We establish the strong consistency and asymptotic normality of the estimator , enabling a computationally efficient, permutation-free Wald test for independence, and an asymptotic confidence interval for the population measure . Extensive simulations and an application to The Cancer Genome Atlas (TCGA) data demonstrate that the proposed method provides coding stability, competitive power, and substantial computational advantages in nominal mixed-type settings.

82 pages, 17 figures. Submitted to the Electronic Journal of Statistics

Topics & keywords

#mixed-type data#association measure#independence testing#nonparametric methods#computational efficiencyξ'ξ_n'label-invariantmonotone transformationWald testasymptotic normalityTCGA
An association measure for mixed-type variables · wovepaper