9 papers
Concentration bounds on response-based vector embeddings of black-box generative models
Aranyak Acharyya, Joshua Agterberg, Youngser Park +1
Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries. Response-based vector embeddings of gener…
Recovering manifold structure in LLM responses through a joint Euclidean mirror
Maximilian Baum, Aranyak Acharyya, Tianyi Chen +5
Understanding the behavior of black-box large language models and determining effective means of comparing their performance is a key task in modern machine learning. We consider h…
Euclidean mirrors and first-order changepoints in network time series
Tianyi Chen, Zachary Lubberts, Avanti Athreya +2
We describe a model for a network time series whose evolution is governed by an underlying stochastic process, known as the latent position process, in which network evolution can…
Multi-rank Subspace Change-point Detection with Application in Monitoring Robotic Swarms
Jonghyeok Lee, Yao Xie, Youngser Park +3
We study real-time detection of low-rank changes in the covariance structure of high-dimensional streaming data, motivated by robotic swarm monitoring. Building on the spiked covar…
Graph Neural Networks Powered by Encoder Embedding for Improved Node Learning
Shiyu Chen, Cencheng Shen, Youngser Park +1
Graph neural networks (GNNs) have emerged as a powerful framework for a wide range of node-level graph learning tasks. However, their performance typically depends on random or min…
Data Kernel Perspective Space Performance Guarantees for Synthetic Data from Transformer Models
Michael Browder, Kevin Duh, J. David Harris +5
Scarcity of labeled training data remains the long pole in the tent for building performant language technology and generative AI models. Transformer models -- particularly LLMs --…