activity
20242026
collaborators

9 papers

stat.ML2026

Concentration bounds on response-based vector embeddings of black-box generative models

Aranyak Acharyya, Joshua Agterberg, Youngser Park +1

Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries. Response-based vector embeddings of gener…

stat.ME2026

Recovering manifold structure in LLM responses through a joint Euclidean mirror

Maximilian Baum, Aranyak Acharyya, Tianyi Chen +5

Understanding the behavior of black-box large language models and determining effective means of comparing their performance is a key task in modern machine learning. We consider h…

stat.ME2026

Euclidean mirrors and first-order changepoints in network time series

Tianyi Chen, Zachary Lubberts, Avanti Athreya +2

We describe a model for a network time series whose evolution is governed by an underlying stochastic process, known as the latent position process, in which network evolution can…

stat.ME2026

Multi-rank Subspace Change-point Detection with Application in Monitoring Robotic Swarms

Jonghyeok Lee, Yao Xie, Youngser Park +3

We study real-time detection of low-rank changes in the covariance structure of high-dimensional streaming data, motivated by robotic swarm monitoring. Building on the spiked covar…

cs.LG2026

Graph Neural Networks Powered by Encoder Embedding for Improved Node Learning

Shiyu Chen, Cencheng Shen, Youngser Park +1

Graph neural networks (GNNs) have emerged as a powerful framework for a wide range of node-level graph learning tasks. However, their performance typically depends on random or min…

cs.CL2026

Data Kernel Perspective Space Performance Guarantees for Synthetic Data from Transformer Models

Michael Browder, Kevin Duh, J. David Harris +5

Scarcity of labeled training data remains the long pole in the tent for building performant language technology and generative AI models. Transformer models -- particularly LLMs --…