Publications (66)
On a phase transition in general order spline regression
Yandi Shen, Qiyang Han, Fang Han
In the Gaussian sequence model in , we study the fundamental limit of approximating the signal by a class of (generalized…
Sterile Neutrino Search Using China Advanced Research Reactor
Gang Guo, Fang Han, Xiangdong Ji +3
We study the feasibility of a sterile neutrino search at the China Advanced Research Reactor by measuring survival probability with a baseline of less than 15 m. Both h…
Bias correction for Chatterjee's graph-based correlation coefficient
Mona Azadkia, Leihao Chen, Fang Han
Azadkia and Chatterjee (2021) recently introduced a simple nearest neighbor (NN) graph-based correlation coefficient that consistently detects both independence and functional depe…
Simultaneous Polysomnography and Cardiotocography Reveal Temporal Correlation Between Maternal Obstructive Sleep Apnea and Fetal Hypoxia
Jingyu Wang, Donglin Xie, Jingying Ma +17
Background: Obstructive sleep apnea syndrome (OSAS) during pregnancy is common and can negatively affect fetal outcomes. However, studies on the immediate effects of maternal hypox…
An Introduction to Permutation Processes (version 0.5)
Fang Han
These lecture notes were prepared for a special topics course in the Department of Statistics at the University of Washington, Seattle. They comprise the first eight chapters of a…
Azadkia-Chatterjee's correlation coefficient adapts to manifold data
Fang Han, Zhihan Huang
In their seminal work, Azadkia and Chatterjee (2021) initiated graph-based methods for measuring variable dependence strength. By appealing to nearest neighbor graphs, they gave an…
Fisher-Pitman permutation tests based on nonparametric Poisson mixtures with application to single cell genomics
Zhen Miao, Weihao Kong, Ramya Korlakai Vinayak +2
This paper investigates the theoretical and empirical performance of Fisher-Pitman-type permutation tests for assessing the equality of unknown Poisson mixture distributions. Build…
Distribution-free tests of multivariate independence based on center-outward quadrant, Spearman, Kendall, and van der Waerden statistics
Hongjian Shi, Mathias Drton, Marc Hallin +1
Due to the lack of a canonical ordering in for , defining multivariate generalizations of the classical univariate ranks has been a long-standing open problem…
On Gaussian Comparison Inequality and Its Application to Spectral Analysis of Large Random Matrices
Fang Han, Sheng Xu, Wen-Xin Zhou
Recently, Chernozhukov, Chetverikov, and Kato [Ann. Statist. 42 (2014) 1564--1597] developed a new Gaussian comparison inequality for approximating the suprema of empirical process…
ECA: High Dimensional Elliptical Component Analysis in non-Gaussian Distributions
Fang Han, Han Liu
We present a robust alternative to principal component analysis (PCA) --- called elliptical component analysis (ECA) --- for analyzing high dimensional, elliptically distributed da…
A Direct Estimation of High Dimensional Stationary Vector Autoregressions
Fang Han, Huanran Lu, Han Liu
The vector autoregressive (VAR) model is a powerful tool in modeling complex time series and has been exploited in many fields. However, fitting high dimensional VAR model poses so…
Neural network an1alysis of sleep stages enables efficient diagnosis of narcolepsy
Jens B. Stephansen, Alexander N. Olesen, Mads Olsen +27
Analysis of sleep for the diagnosis of sleep disorders such as Type-1 Narcolepsy (T1N) currently requires visual inspection of polysomnography records by trained scoring technician…
Decoupling and randomization for double-indexed permutation statistics
Mingxuan Zou, Jingfan Xu, Peng Ding +1
This paper introduces a version of decoupling and randomization to establish concentration inequalities for double-indexed permutation statistics. The results yield, among other ap…
Distribution-Free Tests of Independence in High Dimensions
Fang Han, Shizhe Chen, Han Liu
We consider the testing of mutual independence among all entries in a -dimensional random vector based on independent observations. We study two families of distribution-fre…
Spectral analysis of large dimensional Chatterjee's rank correlation matrix
Zhaorui Dong, Fang Han, Jianfeng Yao
This paper studies the spectral behavior of large dimensional Chatterjee's rank correlation matrix when observations are independent draws from a high-dimensional random vector wit…
A Provable Smoothing Approach for High Dimensional Generalized Regression with Applications in Genomics
Fang Han, Hongkai Ji, Zhicheng Ji +1
In many applications, linear models fit the data poorly. This article studies an appealing alternative, the generalized regression model. This model only assumes that there exists…
Limit theorems of Azadkia-Chatterjee's conditional graph correlation
Muhong Gao, Fang Han, Qizhai Li
Inferring the strength of conditional dependence and testing conditional independence are fundamental problems in statistics. A recent breakthrough by Azadkia and Chatterjee introd…
Limit theorems of Chatterjee's rank correlation
Zhexiao Lin, Fang Han
Establishing the limiting distribution of Chatterjee's rank correlation for a general, possibly non-independent, pair of random variables has been eagerly awaited by many. This pap…
Testing the science/technology relationship by analysis of patent citations of scientific papers after decomposition of both science and technology
Fang Han, Christopher L. Magee
The relationship of scientific knowledge development to technological development is widely recognized as one of the most important and complex aspects of technological evolution.…
A sliced Wasserstein and diffusion approach to random coefficient models
Keunwoo Lim, Ting Ye, Fang Han
We propose a new minimum-distance estimator for linear random coefficient models. This estimator integrates the recently advanced sliced Wasserstein distance with the nearest neigh…
On universally consistent and fully distribution-free rank tests of vector independence
Hongjian Shi, Marc Hallin, Mathias Drton +1
Rank correlations have found many innovative applications in the last decade. In particular, suitable rank correlations have been used for consistent tests of independence between…
High Dimensional Semiparametric Scale-Invariant Principal Component Analysis
Fang Han, Han Liu
We propose a new high dimensional semiparametric principal component analysis (PCA) method, named Copula Component Analysis (COCA). The semiparametric model assumes that, after uns…
Limiting spectral distributions of large consistent rank correlation matrices
Zhaorui Dong, Fang Han, Jianfeng Yao
We study random matrices whose entries are obtained by applying consistent rank correlations, such as Hoeffding's , pairwise to a high-dimensional random vector with mutually in…
Challenges of Big Data Analysis
Jianqing Fan, Fang Han, Han Liu
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heter…
The Nonparanormal SKEPTIC
Han Liu, Fang Han, Ming Yuan +2
We propose a semiparametric approach, named nonparanormal skeptic, for estimating high dimensional undirected graphical models. In terms of modeling, we consider the nonparanormal…
On the failure of the bootstrap for Chatterjee's rank correlation
Zhexiao Lin, Fang Han
While researchers commonly use the bootstrap for statistical inference, many of us have realized that the standard bootstrap, in general, does not work for Chatterjee's rank correl…
STREAMLINE: An Automated Machine Learning Pipeline for Biomedicine Applied to Examine the Utility of Photography-Based Phenotypes for OSA Prediction Across International Sleep Centers
Ryan J. Urbanowicz, Harsh Bandhey, Brendan T. Keenan +19
While machine learning (ML) includes a valuable array of tools for analyzing biomedical data, significant time and expertise is required to assemble effective, rigorous, and unbias…
Bootstrap consistency for general double/debiased machine learning estimators
Ziming Lin, Fang Han
Double/debiased machine learning (DML) provides a general framework for inference with high-dimensional or otherwise complex nuisance parameters by combining Neyman-orthogonal scor…
High dimensional consistent independence testing with maxima of rank correlations
Mathias Drton, Fang Han, Hongjian Shi
Testing mutual independence for high-dimensional observations is a fundamental statistical challenge. Popular tests based on linear and simple rank correlations are known to be inc…
On the power of Chatterjee rank correlation
Hongjian Shi, Mathias Drton, Fang Han
Chatterjee (2021) introduced a simple new rank correlation coefficient that has attracted much recent attention. The coefficient has the unusual appeal that it not only estimates a…
Nonparametric mixture MLEs under Gaussian-smoothed optimal transport distance
Fang Han, Zhen Miao, Yandi Shen
The Gaussian-smoothed optimal transport (GOT) framework, pioneered in Goldfeld et al. (2020) and followed up by a series of subsequent papers, has quickly caught attention among re…
Distribution-free consistent independence tests via center-outward ranks and signs
Hongjian Shi, Mathias Drton, Fang Han
This paper investigates the problem of testing independence of two random vectors of general dimensions. For this, we give for the first time a distribution-free consistent test. O…
Sparse Median Graphs Estimation in a High Dimensional Semiparametric Model
Fang Han, Han Liu, Brian Caffo
In this manuscript a unified framework for conducting inference on complex aggregated data in high dimensional settings is proposed. The data are assumed to be a collection of mult…
Exponential inequalities for dependent V-statistics via random Fourier features
Yandi Shen, Fang Han, Daniela Witten
We establish exponential inequalities for a class of V-statistics under strong mixing conditions. Our theory is developed via a novel kernel expansion based on random Fourier featu…
Generative modeling for the bootstrap
Leon Tran, Ting Ye, Peng Ding +1
Generative modeling builds on and substantially advances the classical idea of simulating synthetic data from observed samples. This paper shows that this principle is not only nat…
High Dimensional Semiparametric Gaussian Copula Graphical Models
Han Liu, Fang Han, Ming Yuan +2
In this paper, we propose a semiparametric approach, named nonparanormal skeptic, for efficiently and robustly estimating high dimensional undirected graphical models. To achieve m…
On the adaptation of causal forests to manifold data
Yiyi Huo, Yingying Fan, Fang Han
Researchers often hold the belief that random forests are "the cure to the world's ills" (Bickel, 2010). But how exactly do they achieve this? Focused on the recently introduced ca…
Probability inequalities for high dimensional time series under a triangular array framework
Fang Han, Weibiao Wu
Study of time series data often involves measuring the strength of temporal dependence, on which statistical properties like consistency and central limit theorem are built. Histor…
Asymptotic joint distribution of extreme eigenvalues and trace of large sample covariance matrix in a generalized spiked population model
Zeng Li, Fang Han, Jianfeng Yao
This paper studies the joint limiting behavior of extreme eigenvalues and trace of large sample covariance matrix in a generalized spiked population model, where the asymptotic reg…
Smoothed NPMLEs in nonparametric Poisson mixtures and beyond
Keunwoo Lim, Fang Han
We discuss nonparametric mixing distribution estimation under the Gaussian-smoothed optimal transport (GOT) distance. It is shown that a recently formulated conjecture -- that the…
Statistical analysis of latent generalized correlation matrix estimation in transelliptical distribution
Fang Han, Han Liu
Correlation matrices play a key role in many multivariate methods (e.g., graphical model estimation and factor analysis). The current state-of-the-art in estimating large correlati…
On Estimation of Isotonic Piecewise Constant Signals
Chao Gao, Fang Han, Cun-Hui Zhang
Consider a sequence of real data points with underlying means . This paper starts from studying the setting that is both piecewise c…
Joint Estimation of Multiple Graphical Models from High Dimensional Time Series
Huitong Qiu, Fang Han, Han Liu +1
In this manuscript we consider the problem of jointly estimating multiple graphical models in high dimensions. We assume that the data are collected from n subjects, each of which…
Robust Inference of Risks of Large Portfolios
Jianqing Fan, Fang Han, Han Liu +1
We propose a bootstrap-based robust high-confidence level upper bound (Robust H-CLUB) for assessing the risks of large portfolios. The proposed approach exploits rank-based and qua…
On propensity score matching with a diverging number of matches
Yihui He, Fang Han
This paper reexamines Abadie and Imbens (2016)'s work on propensity score matching for average treatment effect estimation. We explore the asymptotic behavior of these estimators w…
On Azadkia-Chatterjee's conditional dependence coefficient
Hongjian Shi, Mathias Drton, Fang Han
In recent work, Azadkia and Chatterjee (2021) laid out an ingenious approach to defining consistent measures of conditional dependence. Their fully nonparametric approach forms sta…
An Exponential Inequality for U-Statistics under Mixing Conditions
Fang Han
The family of U-statistics plays a fundamental role in statistics. This paper proves a novel exponential inequality for U-statistics under the time series setting. Explicit mixing…
Sparse Principal Component Analysis for High Dimensional Vector Autoregressive Models
Zhaoran Wang, Fang Han, Han Liu
We study sparse principal component analysis for high dimensional vector autoregressive time series under a doubly asymptotic framework, which allows the dimension to scale wit…
On regression-adjusted imputation estimators of the average treatment effect
Zhexiao Lin, Fang Han
Imputing missing potential outcomes using an estimated regression function is a natural idea for estimating causal effects. In the literature, estimators that combine imputation an…
Estimation based on nearest neighbor matching: from density ratio to average treatment effect
Zhexiao Lin, Peng Ding, Fang Han
Nearest neighbor (NN) matching as a tool to align data sampled from different groups is both conceptually natural and practically well-used. In a landmark paper, Abadie and Imbens…
An Extreme-Value Approach for Testing the Equality of Large U-Statistic Based Correlation Matrices
Cheng Zhou, Fang Han, Xinsheng Zhang +1
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of la…
Robust Functional Principal Component Analysis via Functional Pairwise Spatial Signs
Guangxing Wang, Sisheng Liu, Fang Han +1
Functional principal component analysis (FPCA) has been widely used to capture major modes of variation and reduce dimensions in functional data analysis. However, standard FPCA ba…
On rank estimators in increasing dimensions
Yanqin Fan, Fang Han, Wei Li +1
The family of rank estimators, including Han's maximum rank correlation (Han, 1987) as a notable example, has been widely exploited in studying regression problems. For these estim…
LLM Agents Enable User-Governed Personalization Beyond Platform Boundaries
Jiacheng Lin, Kun Qian, Arvind Srinivasan +15
Personalization today is fundamentally platform-centric: services build user representations from the behavioral fragments they observe. Yet no platform can construct a complete pi…
On boosting the power of Chatterjee's rank correlation
Zhexiao Lin, Fang Han
Chatterjee (2021)'s ingenious approach to estimating a measure of dependence first proposed by Dette et al. (2013) based on simple rank statistics has quickly caught attention. Thi…
Optimal estimation of variance in nonparametric regression with random design
Yandi Shen, Chao Gao, Daniela Witten +1
Consider the heteroscedastic nonparametric regression model with random design \begin{align*} Y_i = f(X_i) + V^{1/2}(X_i)\varepsilon_i, \quad i=1,2,\ldots,n, \end{align*} with $f(\…
Limit theorems of matching estimators with a fixed number of matches
Songliang Chen, Fang Han
This paper re-examines the limit theorems of Abadie and Imbens for nearest-neighbor matching estimators of average treatment effects with a fixed number of matches. We establish, f…
On Rosenbaum's Rank-based Matching Estimator
Matias D. Cattaneo, Fang Han, Zhexiao Lin
In two influential contributions, Rosenbaum (2005, 2020) advocated for using the distances between component-wise ranks, instead of the original data values, to measure covariate s…
Pairwise Difference Estimation of High Dimensional Partially Linear Model
Fang Han, Zhao Ren, Yuxin Zhu
This paper proposes a regularized pairwise difference approach for estimating the linear component coefficient in a partially linear model, with consistency and exact rates of conv…
Tail behavior of dependent V-statistics and its applications
Yandi Shen, Fang Han, Daniela Witten
We establish exponential inequalities and Cramer-type moderate deviation theorems for a class of V-statistics under strong mixing conditions. Our theory is developed via kernel exp…
Moment bounds for large autocovariance matrices under dependence
Fang Han, Yicheng Li
The goal of this paper is to obtain expectation bounds for the deviation of large sample autocovariance matrices from their means under weak data dependence. While the accuracy of…
vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models
Xunzhuo Liu, Huamin Chen, Samzong Lu +30
As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing: selecting the right model for each query a…
On a rank-based Azadkia-Chatterjee correlation coefficient
Leon Tran, Fang Han
Azadkia and Chatterjee (Azadkia and Chatterjee, 2021) recently introduced a graph-based correlation coefficient that has garnered significant attention. The method relies on a near…
On inference validity of weighted U-statistics under data heterogeneity
Fang Han, Tianchen Qian
Motivated by challenges on studying a new correlation measurement being popularized in evaluating online ranking algorithms' performance, this manuscript explores the validity of u…
On the consistency of bootstrap for matching estimators
Ziming Lin, Fang Han
In a landmark paper, Abadie and Imbens (2008) showed that the naive bootstrap is inconsistent when applied to nearest neighbor matching estimators of the average treatment effect w…
On the Impact of Dimension Reduction on Graphical Structures
Fang Han, Huitong Qiu, Han Liu +1
Statisticians and quantitative neuroscientists have actively promoted the use of independence relationships for investigating brain networks, genomic networks, and other measuremen…