papers

Publications (13)

stat.ME2023

Nonlinear Permuted Granger Causality

Noah D. Gade, Jordan Rodu

Granger causal inference is a contentious but widespread method used in fields ranging from economics to neuroscience. The original definition addresses the notion of causality in…

stat.ME2026

Synthetic Data, Information, and Prior Knowledge: Why Synthetic Data Augmentation to Boost Sample Doesn't Work for Statistical Inference

Reid Dale, Jordan Rodu, Mike Baiocchi

The use of synthetic data to deidentify data and to improve predictive models is well-attested to. The augmentation of datasets using synthetically generated data is an alluring pr…

stat.ML2026

Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

Tyler Ashoff, Jordan Rodu

Modern opaque AI models prize performance over interpretability, which makes testing difficult. However, formal statistical tests conducted on a model's embedding space can provide…

stat.AP2015

Locating recombination hot spots in genomic sequences through the singular value decomposition

Jordan Rodu, Shane T. Jensen

Locating recombination hotspots in genomic data is an important but difficult task. Current methods frequently rely on estimating complicated models at high computational cost. In…

cs.CL2021

Trees in transformers: a theoretical analysis of the Transformer's ability to represent trees

Qi He, João Sedoc, Jordan Rodu

Transformer networks are the de facto standard architecture in natural language processing. To date, there are no theoretical analyses of the Transformer's ability to capture tree…

stat.OT2023

When black box algorithms are (not) appropriate: a principled prediction-problem ontology

Jordan Rodu, Michael Baiocchi

In the 1980s a new, extraordinarily productive way of reasoning about algorithms emerged. In this paper, we introduce the term "outcome reasoning" to refer to this form of reasonin…

stat.ML2012

Spectral dimensionality reduction for HMMs

Dean P. Foster, Jordan Rodu, Lyle H. Ungar

Hidden Markov Models (HMMs) can be accurately approximated using co-occurrence frequencies of pairs and triples of observations by using a fast spectral method in contrast to the u…

stat.ML2023

Change Point Detection with Conceptors

Noah D. Gade, Jordan Rodu

Offline change point detection retrospectively locates change points in a time series. Many nonparametric methods that target i.i.d. mean and variance changes fail in the presence…

math.ST2025

Data Gluttony: Epistemic Risks, Dependent Testing and Data Reuse in Large Datasets

Reid Dale, Jordan Rodu, Maria E. Currie +1

Large-scale registries have collected vast amounts of data which has enabled investigators to efficiently conduct studies of observational data. Common practice is for investigator…

math.ST2026

Data Reuse and the Long Shadow of Error: Splitting, Subsampling, and Prospectively Managing Inferential Errors

Reid Dale, Jordan Rodu, Maria E. Currie +1

When multiple investigators analyze a common dataset, the data reuse induces dependence across testing procedures, affecting the distribution of errors. Existing techniques of mana…

stat.ML2024

Bridging the Usability Gap: Theoretical and Methodological Advances for Spectral Learning of Hidden Markov Models

Xiaoyuan Ma, Jordan Rodu

The Baum-Welch (B-W) algorithm is the most widely accepted method for inferring hidden Markov models (HMM). However, it is prone to getting stuck in local optima, and can be too sl…

cs.CL2012

Two Step CCA: A new spectral method for estimating vector models of words

Paramveer Dhillon, Jordan Rodu, Dean Foster +1

Unlabeled data is often used to learn representations which can be used to supplement baseline features in a supervised learner. For example, for text applications where the words…

cs.CL2026

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

Tyler Ashoff, Jordan Rodu

Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly satu…