papers

Publications (66)

math.ST2020

On a phase transition in general order spline regression

Yandi Shen, Qiyang Han, Fang Han

In the Gaussian sequence model in , we study the fundamental limit of approximating the signal by a class of (generalized…

physics.ins-det2013

Sterile Neutrino Search Using China Advanced Research Reactor

Gang Guo, Fang Han, Xiangdong Ji +3

We study the feasibility of a sterile neutrino search at the China Advanced Research Reactor by measuring survival probability with a baseline of less than 15 m. Both h…

stat.ME2026

Bias correction for Chatterjee's graph-based correlation coefficient

Mona Azadkia, Leihao Chen, Fang Han

Azadkia and Chatterjee (2021) recently introduced a simple nearest neighbor (NN) graph-based correlation coefficient that consistently detects both independence and functional depe…

eess.SP2025

Simultaneous Polysomnography and Cardiotocography Reveal Temporal Correlation Between Maternal Obstructive Sleep Apnea and Fetal Hypoxia

Jingyu Wang, Donglin Xie, Jingying Ma +17

Background: Obstructive sleep apnea syndrome (OSAS) during pregnancy is common and can negatively affect fetal outcomes. However, studies on the immediate effects of maternal hypox…

math.ST2024

An Introduction to Permutation Processes (version 0.5)

Fang Han

These lecture notes were prepared for a special topics course in the Department of Statistics at the University of Washington, Seattle. They comprise the first eight chapters of a…

math.ST2022

Azadkia-Chatterjee's correlation coefficient adapts to manifold data

Fang Han, Zhihan Huang

In their seminal work, Azadkia and Chatterjee (2021) initiated graph-based methods for measuring variable dependence strength. By appealing to nearest neighbor graphs, they gave an…

stat.ME2021

Fisher-Pitman permutation tests based on nonparametric Poisson mixtures with application to single cell genomics

Zhen Miao, Weihao Kong, Ramya Korlakai Vinayak +2

This paper investigates the theoretical and empirical performance of Fisher-Pitman-type permutation tests for assessing the equality of unknown Poisson mixture distributions. Build…

math.ST2024

Distribution-free tests of multivariate independence based on center-outward quadrant, Spearman, Kendall, and van der Waerden statistics

Hongjian Shi, Mathias Drton, Marc Hallin +1

Due to the lack of a canonical ordering in for , defining multivariate generalizations of the classical univariate ranks has been a long-standing open problem…

math.ST2017

On Gaussian Comparison Inequality and Its Application to Spectral Analysis of Large Random Matrices

Fang Han, Sheng Xu, Wen-Xin Zhou

Recently, Chernozhukov, Chetverikov, and Kato [Ann. Statist. 42 (2014) 1564--1597] developed a new Gaussian comparison inequality for approximating the suprema of empirical process…

stat.ML2016

ECA: High Dimensional Elliptical Component Analysis in non-Gaussian Distributions

Fang Han, Han Liu

We present a robust alternative to principal component analysis (PCA) --- called elliptical component analysis (ECA) --- for analyzing high dimensional, elliptically distributed da…

stat.ML2014

A Direct Estimation of High Dimensional Stationary Vector Autoregressions

Fang Han, Huanran Lu, Han Liu

The vector autoregressive (VAR) model is a powerful tool in modeling complex time series and has been exploited in many fields. However, fitting high dimensional VAR model poses so…

cs.NE2019

Neural network an1alysis of sleep stages enables efficient diagnosis of narcolepsy

Jens B. Stephansen, Alexander N. Olesen, Mads Olsen +27

Analysis of sleep for the diagnosis of sleep disorders such as Type-1 Narcolepsy (T1N) currently requires visual inspection of polysomnography records by trained scoring technician…

math.ST2026

Decoupling and randomization for double-indexed permutation statistics

Mingxuan Zou, Jingfan Xu, Peng Ding +1

This paper introduces a version of decoupling and randomization to establish concentration inequalities for double-indexed permutation statistics. The results yield, among other ap…

math.ST2017

Distribution-Free Tests of Independence in High Dimensions

Fang Han, Shizhe Chen, Han Liu

We consider the testing of mutual independence among all entries in a -dimensional random vector based on independent observations. We study two families of distribution-fre…

math.ST2026

Spectral analysis of large dimensional Chatterjee's rank correlation matrix

Zhaorui Dong, Fang Han, Jianfeng Yao

This paper studies the spectral behavior of large dimensional Chatterjee's rank correlation matrix when observations are independent draws from a high-dimensional random vector wit…

stat.ME2017

A Provable Smoothing Approach for High Dimensional Generalized Regression with Applications in Genomics

Fang Han, Hongkai Ji, Zhicheng Ji +1

In many applications, linear models fit the data poorly. This article studies an appealing alternative, the generalized regression model. This model only assumes that there exists…

math.ST2026

Limit theorems of Azadkia-Chatterjee's conditional graph correlation

Muhong Gao, Fang Han, Qizhai Li

Inferring the strength of conditional dependence and testing conditional independence are fundamental problems in statistics. A recent breakthrough by Azadkia and Chatterjee introd…

math.ST2025

Limit theorems of Chatterjee's rank correlation

Zhexiao Lin, Fang Han

Establishing the limiting distribution of Chatterjee's rank correlation for a general, possibly non-independent, pair of random variables has been eagerly awaited by many. This pap…

cs.DL2017

Testing the science/technology relationship by analysis of patent citations of scientific papers after decomposition of both science and technology

Fang Han, Christopher L. Magee

The relationship of scientific knowledge development to technological development is widely recognized as one of the most important and complex aspects of technological evolution.…

math.ST2025

A sliced Wasserstein and diffusion approach to random coefficient models

Keunwoo Lim, Ting Ye, Fang Han

We propose a new minimum-distance estimator for linear random coefficient models. This estimator integrates the recently advanced sliced Wasserstein distance with the nearest neigh…

math.ST2021

On universally consistent and fully distribution-free rank tests of vector independence

Hongjian Shi, Marc Hallin, Mathias Drton +1

Rank correlations have found many innovative applications in the last decade. In particular, suitable rank correlations have been used for consistent tests of independence between…

stat.ML2014

High Dimensional Semiparametric Scale-Invariant Principal Component Analysis

Fang Han, Han Liu

We propose a new high dimensional semiparametric principal component analysis (PCA) method, named Copula Component Analysis (COCA). The semiparametric model assumes that, after uns…

math.PR2026

Limiting spectral distributions of large consistent rank correlation matrices

Zhaorui Dong, Fang Han, Jianfeng Yao

We study random matrices whose entries are obtained by applying consistent rank correlations, such as Hoeffding's , pairwise to a high-dimensional random vector with mutually in…

stat.ML2014

Challenges of Big Data Analysis

Jianqing Fan, Fang Han, Han Liu

Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heter…

stat.ME2012

The Nonparanormal SKEPTIC

Han Liu, Fang Han, Ming Yuan +2

We propose a semiparametric approach, named nonparanormal skeptic, for estimating high dimensional undirected graphical models. In terms of modeling, we consider the nonparanormal…

math.ST2023

On the failure of the bootstrap for Chatterjee's rank correlation

Zhexiao Lin, Fang Han

While researchers commonly use the bootstrap for statistical inference, many of us have realized that the standard bootstrap, in general, does not work for Chatterjee's rank correl…

cs.LG2023

STREAMLINE: An Automated Machine Learning Pipeline for Biomedicine Applied to Examine the Utility of Photography-Based Phenotypes for OSA Prediction Across International Sleep Centers

Ryan J. Urbanowicz, Harsh Bandhey, Brendan T. Keenan +19

While machine learning (ML) includes a valuable array of tools for analyzing biomedical data, significant time and expertise is required to assemble effective, rigorous, and unbias…

math.ST2026

Bootstrap consistency for general double/debiased machine learning estimators

Ziming Lin, Fang Han

Double/debiased machine learning (DML) provides a general framework for inference with high-dimensional or otherwise complex nuisance parameters by combining Neyman-orthogonal scor…

math.ST2020

High dimensional consistent independence testing with maxima of rank correlations

Mathias Drton, Fang Han, Hongjian Shi

Testing mutual independence for high-dimensional observations is a fundamental statistical challenge. Popular tests based on linear and simple rank correlations are known to be inc…

math.ST2021

On the power of Chatterjee rank correlation

Hongjian Shi, Mathias Drton, Fang Han

Chatterjee (2021) introduced a simple new rank correlation coefficient that has attracted much recent attention. The coefficient has the unusual appeal that it not only estimates a…

math.ST2021

Nonparametric mixture MLEs under Gaussian-smoothed optimal transport distance

Fang Han, Zhen Miao, Yandi Shen

The Gaussian-smoothed optimal transport (GOT) framework, pioneered in Goldfeld et al. (2020) and followed up by a series of subsequent papers, has quickly caught attention among re…

math.ST2020

Distribution-free consistent independence tests via center-outward ranks and signs

Hongjian Shi, Mathias Drton, Fang Han

This paper investigates the problem of testing independence of two random vectors of general dimensions. For this, we give for the first time a distribution-free consistent test. O…

stat.AP2013

Sparse Median Graphs Estimation in a High Dimensional Semiparametric Model

Fang Han, Han Liu, Brian Caffo

In this manuscript a unified framework for conducting inference on complex aggregated data in high dimensional settings is proposed. The data are assumed to be a collection of mult…

math.ST2020

Exponential inequalities for dependent V-statistics via random Fourier features

Yandi Shen, Fang Han, Daniela Witten

We establish exponential inequalities for a class of V-statistics under strong mixing conditions. Our theory is developed via a novel kernel expansion based on random Fourier featu…

stat.ME2026

Generative modeling for the bootstrap

Leon Tran, Ting Ye, Peng Ding +1

Generative modeling builds on and substantially advances the classical idea of simulating synthetic data from observed samples. This paper shows that this principle is not only nat…

stat.ML2012

High Dimensional Semiparametric Gaussian Copula Graphical Models

Han Liu, Fang Han, Ming Yuan +2

In this paper, we propose a semiparametric approach, named nonparanormal skeptic, for efficiently and robustly estimating high dimensional undirected graphical models. To achieve m…

math.ST2023

On the adaptation of causal forests to manifold data

Yiyi Huo, Yingying Fan, Fang Han

Researchers often hold the belief that random forests are "the cure to the world's ills" (Bickel, 2010). But how exactly do they achieve this? Focused on the recently introduced ca…

math.ST2019

Probability inequalities for high dimensional time series under a triangular array framework

Fang Han, Weibiao Wu

Study of time series data often involves measuring the strength of temporal dependence, on which statistical properties like consistency and central limit theorem are built. Histor…

math.ST2019

Asymptotic joint distribution of extreme eigenvalues and trace of large sample covariance matrix in a generalized spiked population model

Zeng Li, Fang Han, Jianfeng Yao

This paper studies the joint limiting behavior of extreme eigenvalues and trace of large sample covariance matrix in a generalized spiked population model, where the asymptotic reg…

math.ST2024

Smoothed NPMLEs in nonparametric Poisson mixtures and beyond

Keunwoo Lim, Fang Han

We discuss nonparametric mixing distribution estimation under the Gaussian-smoothed optimal transport (GOT) distance. It is shown that a recently formulated conjecture -- that the…

stat.ML2016

Statistical analysis of latent generalized correlation matrix estimation in transelliptical distribution

Fang Han, Han Liu

Correlation matrices play a key role in many multivariate methods (e.g., graphical model estimation and factor analysis). The current state-of-the-art in estimating large correlati…

math.ST2019

On Estimation of Isotonic Piecewise Constant Signals

Chao Gao, Fang Han, Cun-Hui Zhang

Consider a sequence of real data points with underlying means . This paper starts from studying the setting that is both piecewise c…

stat.ML2014

Joint Estimation of Multiple Graphical Models from High Dimensional Time Series

Huitong Qiu, Fang Han, Han Liu +1

In this manuscript we consider the problem of jointly estimating multiple graphical models in high dimensions. We assume that the data are collected from n subjects, each of which…

math.ST2015

Robust Inference of Risks of Large Portfolios

Jianqing Fan, Fang Han, Han Liu +1

We propose a bootstrap-based robust high-confidence level upper bound (Robust H-CLUB) for assessing the risks of large portfolios. The proposed approach exploits rank-based and qua…

math.ST2023

On propensity score matching with a diverging number of matches

Yihui He, Fang Han

This paper reexamines Abadie and Imbens (2016)'s work on propensity score matching for average treatment effect estimation. We explore the asymptotic behavior of these estimators w…

math.ST2022

On Azadkia-Chatterjee's conditional dependence coefficient

Hongjian Shi, Mathias Drton, Fang Han

In recent work, Azadkia and Chatterjee (2021) laid out an ingenious approach to defining consistent measures of conditional dependence. Their fully nonparametric approach forms sta…

math.ST2016

An Exponential Inequality for U-Statistics under Mixing Conditions

Fang Han

The family of U-statistics plays a fundamental role in statistics. This paper proves a novel exponential inequality for U-statistics under the time series setting. Explicit mixing…

stat.ML2013

Sparse Principal Component Analysis for High Dimensional Vector Autoregressive Models

Zhaoran Wang, Fang Han, Han Liu

We study sparse principal component analysis for high dimensional vector autoregressive time series under a doubly asymptotic framework, which allows the dimension to scale wit…

math.ST2023

On regression-adjusted imputation estimators of the average treatment effect

Zhexiao Lin, Fang Han

Imputing missing potential outcomes using an estimated regression function is a natural idea for estimating causal effects. In the literature, estimators that combine imputation an…

math.ST2021

Estimation based on nearest neighbor matching: from density ratio to average treatment effect

Zhexiao Lin, Peng Ding, Fang Han

Nearest neighbor (NN) matching as a tool to align data sampled from different groups is both conceptually natural and practically well-used. In a landmark paper, Abadie and Imbens…

math.ST2018

An Extreme-Value Approach for Testing the Equality of Large U-Statistic Based Correlation Matrices

Cheng Zhou, Fang Han, Xinsheng Zhang +1

There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of la…

stat.ME2021

Robust Functional Principal Component Analysis via Functional Pairwise Spatial Signs

Guangxing Wang, Sisheng Liu, Fang Han +1

Functional principal component analysis (FPCA) has been widely used to capture major modes of variation and reduce dimensions in functional data analysis. However, standard FPCA ba…

math.ST2019

On rank estimators in increasing dimensions

Yanqin Fan, Fang Han, Wei Li +1

The family of rank estimators, including Han's maximum rank correlation (Han, 1987) as a notable example, has been widely exploited in studying regression problems. For these estim…

cs.IR2026

LLM Agents Enable User-Governed Personalization Beyond Platform Boundaries

Jiacheng Lin, Kun Qian, Arvind Srinivasan +15

Personalization today is fundamentally platform-centric: services build user representations from the behavioral fragments they observe. Yet no platform can construct a complete pi…

math.ST2021

On boosting the power of Chatterjee's rank correlation

Zhexiao Lin, Fang Han

Chatterjee (2021)'s ingenious approach to estimating a measure of dependence first proposed by Dette et al. (2013) based on simple rank statistics has quickly caught attention. Thi…

math.ST2020

Optimal estimation of variance in nonparametric regression with random design

Yandi Shen, Chao Gao, Daniela Witten +1

Consider the heteroscedastic nonparametric regression model with random design \begin{align*} Y_i = f(X_i) + V^{1/2}(X_i)\varepsilon_i, \quad i=1,2,\ldots,n, \end{align*} with $f(\…

math.ST2026

Limit theorems of matching estimators with a fixed number of matches

Songliang Chen, Fang Han

This paper re-examines the limit theorems of Abadie and Imbens for nearest-neighbor matching estimators of average treatment effects with a fixed number of matches. We establish, f…

math.ST2024

On Rosenbaum's Rank-based Matching Estimator

Matias D. Cattaneo, Fang Han, Zhexiao Lin

In two influential contributions, Rosenbaum (2005, 2020) advocated for using the distances between component-wise ranks, instead of the original data values, to measure covariate s…

math.ST2018

Pairwise Difference Estimation of High Dimensional Partially Linear Model

Fang Han, Zhao Ren, Yuxin Zhu

This paper proposes a regularized pairwise difference approach for estimating the linear component coefficient in a partially linear model, with consistency and exact rates of conv…

math.ST2019

Tail behavior of dependent V-statistics and its applications

Yandi Shen, Fang Han, Daniela Witten

We establish exponential inequalities and Cramer-type moderate deviation theorems for a class of V-statistics under strong mixing conditions. Our theory is developed via kernel exp…

math.ST2019

Moment bounds for large autocovariance matrices under dependence

Fang Han, Yicheng Li

The goal of this paper is to obtain expectation bounds for the deviation of large sample autocovariance matrices from their means under weak data dependence. While the accuracy of…

cs.NI2026

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

Xunzhuo Liu, Huamin Chen, Samzong Lu +30

As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing: selecting the right model for each query a…

math.ST2024

On a rank-based Azadkia-Chatterjee correlation coefficient

Leon Tran, Fang Han

Azadkia and Chatterjee (Azadkia and Chatterjee, 2021) recently introduced a graph-based correlation coefficient that has garnered significant attention. The method relies on a near…

math.ST2018

On inference validity of weighted U-statistics under data heterogeneity

Fang Han, Tianchen Qian

Motivated by challenges on studying a new correlation measurement being popularized in evaluating online ranking algorithms' performance, this manuscript explores the validity of u…

math.ST2024

On the consistency of bootstrap for matching estimators

Ziming Lin, Fang Han

In a landmark paper, Abadie and Imbens (2008) showed that the naive bootstrap is inconsistent when applied to nearest neighbor matching estimators of the average treatment effect w…

stat.ME2014

On the Impact of Dimension Reduction on Graphical Structures

Fang Han, Huitong Qiu, Han Liu +1

Statisticians and quantitative neuroscientists have actively promoted the use of independence relationships for investigating brain networks, genomic networks, and other measuremen…