papers

Publications (55)

stat.ME2019

Central Limit Theorems for Classical Multidimensional Scaling

Gongkai Li, Minh Tang, Nichlas Charon +1

Classical multidimensional scaling is a widely used method in dimensionality reduction and manifold learning. The method takes in a dissimilarity matrix and outputs a low-dimension…

stat.ME2021

Valid Two-Sample Graph Testing via Optimal Transport Procrustes and Multiscale Graph Correlation with Applications in Connectomics

Jaewon Chung, Bijan Varjavand, Jesus Arroyo +5

Testing whether two graphs come from the same distribution is of interest in many real world scenarios, including brain network analysis. Under the random dot product graph model,…

math.ST2018

On spectral embedding performance and elucidating network structure in stochastic block model graphs

Joshua Cape, Minh Tang, Carey E. Priebe

Statistical inference on graphs often proceeds via spectral methods involving low-dimensional embeddings of matrix-valued graph representations, such as the graph Laplacian or adja…

math.ST2018

The two-to-infinity norm and singular subspace geometry with applications to high-dimensional statistics

Joshua Cape, Minh Tang, Carey E. Priebe

The singular value matrix decomposition plays a ubiquitous role throughout statistics and related fields. Myriad applications including clustering, classification, and dimensionali…

stat.ML2013

On latent position inference from doubly stochastic messaging activities

Nam H. Lee, Jordan Yoder, Minh Tang +1

We model messaging activities as a hierarchical doubly stochastic point process with three main levels, and develop an iterative algorithm for inferring actors' relative latent pos…

stat.ML2021

A statistical interpretation of spectral embedding: the generalised random dot product graph

Patrick Rubin-Delanchy, Joshua Cape, Minh Tang +1

Spectral embedding is a procedure which can be used to obtain vector representations of the nodes of a graph. This paper proposes a generalisation of the latent position network mo…

stat.ML2012

Generalized Canonical Correlation Analysis for Disparate Data Fusion

Ming Sun, Carey E. Priebe, Minh Tang

Manifold matching works to identify embeddings of multiple disparate data spaces into the same low-dimensional space, where joint inference can be pursued. It is an enabling method…

stat.ME2020

Two-sample Testing on Latent Distance Graphs With Unknown Link Functions

Yiran Wang, Minh Tang, Soumendra Nath Lahiri

We propose a valid and consistent test for the hypothesis that two latent distance random graphs on the same vertex set have the same generating latent positions, up to some uniden…

stat.ME2012

Consistent adjacency-spectral partitioning for the stochastic block model when the model parameters are unknown

Donniell E. Fishkind, Daniel L. Sussman, Minh Tang +2

For random graphs distributed according to a stochastic block model, we consider the inferential task of partioning vertices into blocks using spectral techniques. Spectral partion…

stat.ML2022

Exact Recovery of Community Structures Using DeepWalk and Node2vec

Yichi Zhang, Minh Tang

Random-walk based network embedding algorithms like DeepWalk and node2vec are widely used to obtain Euclidean representation of the nodes in a network prior to performing downstrea…

stat.ML2017

Semiparametric spectral modeling of the Drosophila connectome

Carey E. Priebe, Youngser Park, Minh Tang +8

We present semiparametric spectral modeling of the complete larval Drosophila mushroom body connectome. Motivated by a thorough exploratory data analysis of the network via Gaussia…

stat.ML2012

A consistent adjacency spectral embedding for stochastic blockmodel graphs

Daniel L. Sussman, Minh Tang, Donniell E. Fishkind +1

We present a method to estimate block membership of nodes in a random graph generated by a stochastic blockmodel. We use an embedding procedure motivated by the random dot product…

stat.ME2017

Asymptotically efficient estimators for stochastic blockmodels: the naive MLE, the rank-constrained MLE, and the spectral

Minh Tang, Joshua Cape, Carey E. Priebe

We establish asymptotic normality results for estimation of the block probability matrix in stochastic blockmodel graphs using spectral embedding when the average degr…

stat.ML2024

Regression for matrix-valued data via Kronecker products factorization

Yin-Jen Chen, Minh Tang

We study the matrix-variate regression problem $Y_i = \sum_{k} β_{1k} X_i β_{2k}^{\top} + E_i$ for in the high dimensional regime wherein the response are ma…

stat.ML2016

Community Detection and Classification in Hierarchical Stochastic Blockmodels

Vince Lyzinski, Minh Tang, Avanti Athreya +2

We propose a robust, scalable, integrated methodology for community detection and community comparison in graphs. In our procedure, we first embed a graph into an appropriate Eucli…

stat.ML2019

On a 'Two Truths' Phenomenon in Spectral Graph Clustering

Carey E. Priebe, Youngser Park, Joshua T. Vogelstein +6

Clustering is concerned with coherently grouping observations without any explicit concept of true groupings. Spectral graph clustering - clustering the vertices of a graph based o…

stat.ME2020

On estimation and inference in latent structure random graphs

Avanti Athreya, Minh Tang, Youngser Park +1

We define a latent structure model (LSM) random graph as a random dot product graph (RDPG) in which the latent position distribution incorporates both probabilistic and geometric c…

stat.ME2025

Chain-linked multiple matrix integration via embedding alignment

Runbing Zheng, Minh Tang

Motivated by the increasing demand for multi-source data integration in various scientific fields, in this paper we study matrix completion in scenarios where the data exhibits cer…

math.ST2020

On Two Distinct Sources of Nonidentifiability in Latent Position Random Graph Models

Joshua Agterberg, Minh Tang, Carey E. Priebe

Two separate and distinct sources of nonidentifiability arise naturally in the context of latent position random graph models, though neither are unique to this setting. In this pa…

stat.ME2023

Independence testing for inhomogeneous random graphs

Yukun Song, Carey E. Priebe, Minh Tang

Testing for independence between graphs is a problem that arises naturally in social network analysis and neuroscience. In this paper, we address independence testing for inhomogen…

stat.ME2026

Predictive Subsampling for Scalable Inference in Networks

Arpan Kumar, Minh Tang, Srijan Sengupta

Current methods for statistical inference in networks often encounter substantial computational bottlenecks when applied to the massive network datasets that are increasingly commo…

stat.ME2016

Empirical Bayes Estimation for the Stochastic Blockmodel

Shakira Suwan, Dominic S. Lee, Runze Tang +3

Inference for the stochastic blockmodel is currently of burgeoning interest in the statistical community, as well as in various application domains as diverse as social networks, c…

stat.ML2013

Universally consistent vertex classification for latent positions graphs

Minh Tang, Daniel L. Sussman, Carey E. Priebe

In this work we show that, using the eigen-decomposition of the adjacency matrix, we can consistently estimate feature maps for latent position graphs with positive definite link f…

stat.ML2021

Learning 1-Dimensional Submanifolds for Subsequent Inference on Random Dot Product Graphs

Michael W. Trosset, Mingyue Gao, Minh Tang +1

A random dot product graph (RDPG) is a generative model for networks in which vertices correspond to positions in a latent Euclidean space and edge probabilities are determined by…

math.ST2025

Limit results for distributed estimation of invariant subspaces in multiple networks inference and PCA

Runbing Zheng, Minh Tang

Several statistical problems, such as multiple heterogeneous graph analysis, distributed PCA, integrative data analysis, and simultaneous dimension reduction of images, can involve…

stat.ML2014

Statistical inference on errorfully observed graphs

Carey E. Priebe, Daniel L. Sussman, Minh Tang +1

Statistical inference on graphs is a burgeoning field in the applied and theoretical statistics communities, as well as throughout the wider world of science, engineering, business…

stat.ML2012

Universally Consistent Latent Position Estimation and Vertex Classification for Random Dot Product Graphs

Daniel L. Sussman, Minh Tang, Carey E. Priebe

In this work we show that, using the eigen-decomposition of the adjacency matrix, we can consistently estimate latent positions for random dot product graphs provided the latent po…

stat.ML2014

Generalized Canonical Correlation Analysis for Classification

Cencheng Shen, Ming Sun, Minh Tang +1

For multiple multivariate data sets, we derive conditions under which Generalized Canonical Correlation Analysis (GCCA) improves classification performance of the projected dataset…

stat.ME2025

A Unified Framework for Community Detection and Model Selection in Blockmodels

Subhankar Bhadra, Minh Tang, Srijan Sengupta

Blockmodels are a foundational tool for modeling community structure in networks, with the stochastic blockmodel (SBM), degree-corrected blockmodel (DCBM), and popularity-adjusted…

math.ST2015

A nonparametric two-sample hypothesis testing problem for random dot product graphs

Minh Tang, Avanti Athreya, Daniel L. Sussman +2

We consider the problem of testing whether two finite-dimensional random dot product graphs have generating latent positions that are independently drawn from the same distribution…

math.ST2026

Nonparametric two-sample hypothesis testing for low-rank random graphs of differing sizes

Joshua Agterberg, Minh Tang, Carey Priebe

Given two networks of differing sizes, it is of interest to test whether the two networks belong to the same distribution. We formalize the notion of "equality of distribution" und…

stat.ML2021

Supervised Dimensionality Reduction for Big Data

Joshua T. Vogelstein, Eric Bridgeford, Minh Tang +4

To solve key biomedical problems, experimentalists now routinely measure millions or billions of features (dimensions) per sample, with the hope that data science techniques will b…

math.ST2025

Perturbation Analysis of Randomized SVD and its Applications to Statistics

Yichi Zhang, Minh Tang

Randomized singular value decomposition (RSVD) is a class of computationally efficient algorithms for computing the truncated SVD of large data matrices. Given an matr…

stat.ML2015

Perfect Clustering for Stochastic Blockmodel Graphs via Adjacency Spectral Embedding

Vince Lyzinski, Daniel Sussman, Minh Tang +2

Vertex clustering in a stochastic blockmodel graph has wide applicability and has been the subject of extensive research. In thispaper, we provide a short proof that the adjacency…

math.ST2013

A central limit theorem for scaled eigenvectors of random dot product graphs

Avanti Athreya, Vince Lyzinski, David J. Marchette +3

We prove a central limit theorem for the components of the largest eigenvectors of the adjacency matrix of a finite-dimensional random dot product graph whose true latent positions…

stat.ME2017

Statistical inference on random dot product graphs: a survey

Avanti Athreya, Donniell E. Fishkind, Keith Levin +7

The random dot product graph (RDPG) is an independent-edge random graph that is analytically tractable and, simultaneously, either encompasses or can successfully approximate a wid…

stat.CO2020

Numerical tolerance for spectral decompositions of random matrices

Avanti Athreya, Michael Kane, Bryan Lewis +5

We precisely quantify the impact of statistical error in the quality of a numerical approximation to a random matrix eigendecomposition, and under mild conditions, we use this to i…

math.ST2018

Signal-plus-noise matrix models: eigenvector deviations and fluctuations

Joshua Cape, Minh Tang, Carey E. Priebe

Estimating eigenvectors and low-dimensional subspaces is of central importance for numerous problems in statistics, computer science, and applied mathematics. This paper characteri…

stat.ME2017

Robust Estimation from Multiple Graphs under Gross Error Contamination

Runze Tang, Minh Tang, Joshua T. Vogelstein +1

Estimation of graph parameters based on a collection of graphs is essential for a wide range of graph inference tasks. In practice, weighted graphs are generally observed with edge…

stat.ME2017

Consistency of adjacency spectral embedding for the mixed membership stochastic blockmodel

Patrick Rubin-Delanchy, Carey E. Priebe, Minh Tang

The mixed membership stochastic blockmodel is a statistical model for a graph, which extends the stochastic blockmodel by allowing every node to randomly choose a different communi…

stat.ML2026

Classification of high-dimensional data with spiked covariance matrix structure

Yin-Jen Chen, Minh Tang

We study the classification problem for high-dimensional data with observations on features where the covariance matrix exhibits a spiked eigenvalue struc…

stat.ML2018

The eigenvalues of stochastic blockmodel graphs

Minh Tang

We derive the limiting distribution for the largest eigenvalues of the adjacency matrix for a stochastic blockmodel graph when the number of vertices tends to infinity. We show tha…

stat.AP2013

Locality statistics for anomaly detection in time series of graphs

Heng Wang, Minh Tang, Youngser Park +1

The ability to detect change-points in a dynamic network or a time series of graphs is an increasingly important task in many applications of the emerging discipline of graph signa…

stat.ME2022

Hypothesis Testing for Equality of Latent Positions in Random Graphs

Xinjie Du, Minh Tang

We consider the hypothesis testing problem that two vertices and of a generalized random dot product graph have the same latent positions, possibly up to scaling. Special c…

stat.ML2019

Limit theorems for out-of-sample extensions of the adjacency and Laplacian spectral embeddings

Keith Levin, Fred Roosta, Minh Tang +2

Graph embeddings, a class of dimensionality reduction techniques designed for relational data, have proven useful in exploring and modeling network structure. Most dimensionality r…

stat.ML2022

Adversarial contamination of networks in the setting of vertex nomination: a new trimming method

Sheyda Peyman, Minh Tang, Vince Lyzinski

As graph data becomes more ubiquitous, the need for robust inferential graph algorithms to operate in these complex data domains is crucial. In many cases of interest, inference is…

stat.ML2013

Out-of-sample Extension for Latent Position Graphs

Minh Tang, Youngser Park, Carey E. Priebe

We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some…

stat.ME2015

A semiparametric two-sample hypothesis testing problem for random dot product graphs

Minh Tang, Avanti Athreya, Daniel L. Sussman +2

Two-sample hypothesis testing for random graphs arises naturally in neuroscience, social networks, and machine learning. In this paper, we consider a semiparametric problem of two-…

stat.ML2025

Out-of-Sample Embedding with Proximity Data: Projection versus Restricted Reconstruction

Michael W. Trosset, Kaiyi Tan, Minh Tang +1

The problem of using proximity (similarity or dissimilarity) data for the purpose of "adding a point to a vector diagram" was first studied by J.C. Gower in 1968. Since then, a num…

stat.ML2022

Popularity Adjusted Block Models are Generalized Random Dot Product Graphs

John Koo, Minh Tang, Michael W. Trosset

We connect two random graph models, the Popularity Adjusted Block Model (PABM) and the Generalized Random Dot Product Graph (GRDPG), by demonstrating that the PABM is a special cas…

cs.SI2022

Vertex nomination between graphs via spectral embedding and quadratic programming

Runbing Zheng, Vince Lyzinski, Carey E. Priebe +1

Given a network and a subset of interesting vertices whose identities are only partially known, the vertex nomination problem seeks to rank the remaining vertices in such a way tha…

stat.ML2016

Limit theorems for eigenvectors of the normalized Laplacian for random graphs

Minh Tang, Carey E. Priebe

We prove a central limit theorem for the components of the eigenvectors corresponding to the largest eigenvalues of the normalized Laplacian matrix of a finite dimensional rand…

stat.ME2019

A central limit theorem for an omnibus embedding of multiple random graphs and implications for multiscale network inference

Keith Levin, Avanti Athreya, Minh Tang +3

Performing statistical analyses on collections of graphs is of import to many disciplines, but principled, scalable methods for multi-sample graph inference are few. Here we descri…

math.ST2018

The Kato--Temple inequality and eigenvalue concentration with applications to graph inference

Joshua Cape, Minh Tang, Carey E. Priebe

We present an adaptation of the Kato--Temple inequality for bounding perturbations of eigenvalues with applications to statistical inference for random graphs, specifically hypothe…

math.ST2025

Eigenvector fluctuations and limit results for random graphs with infinite rank kernels

Minh Tang, Joshua R. Cape

This paper systematically studies the behavior of the leading eigenvectors for independent edge undirected random graphs generated from a general latent position model whose link f…