8 papers
A Theory on Flow Matching with Neural Networks
Yihan He, Qishuo Yin, Yuan Cao +2
In this work, we develop theoretical foundation for flow matching with neural-network-parameterized conditional velocity fields. We establish convergence guarantees for gradient de…
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
Yuan Cao, Mingyang Wang, Hinrich Schütze
Large language models (LLMs) are increasingly used as knowledge bases, but keeping them up to date requires targeted knowledge editing (KE). However, it remains unclear how edits a…
Transformers Simulate MLE for Sequence Generation in Bayesian Networks
Yuan Cao, Yihan He, Dennis Wu +3
Transformers have achieved significant success in various fields, notably excelling in tasks involving sequential data like natural language processing. Despite these achievements,…
Transformers versus the EM Algorithm in Multi-class Clustering
Yihan He, Hong-Yu Chen, Yuan Cao +2
LLMs demonstrate significant inference capacities in complicated machine learning tasks, using the Transformer model as its backbone. Motivated by the limited understanding of such…
Transformers and Their Roles as Time Series Foundation Models
Dennis Wu, Yihan He, Yuan Cao +2
We give a comprehensive analysis of transformers as time series foundation models, focusing on their approximation and generalization capabilities. First, we demonstrate that there…
Learning Spectral Methods by Transformers
Yihan He, Yuan Cao, Hong-Yu Chen +3
Transformers demonstrate significant advantages as the building block of modern LLMs. In this work, we study the capacities of Transformers in performing unsupervised learning. We…