collaborators

7 papers

cs.LG2026

Transformer-like Inference from Optimal Control

Aditya Kudre, Heng-Sheng Chang, Prashant G. Mehta

Decoder-only transformers compute the conditional probability of the next token from a sequence of past observations. This paper derives, from first principles, inference architect…

cs.LG2026

Differentiable Filtering for Learning Hidden Markov Models

Reginald Zhiyan Chen, Heng-Sheng Chang, Prashant G. Mehta

Hidden Markov Models (HMMs) are fundamental for modeling sequential data, yet learning their parameters from observations remains challenging. Classical methods like the Baum-Welch…

eess.SY2026

Duality Theory for Non-Markovian Linear Gaussian Models

Aditya Kudre, Heng-Sheng Chang, Prashant G. Mehta

This work develops a duality theory for partially observed linear Gaussian models in discrete time. The state process evolves according to a causal but non-Markovian (or higher-ord…

cs.LG2026

Dual Filter: A Transformer-like Inference Architecture for Hidden Markov Models

Heng-Sheng Chang, Prashant G. Mehta

This paper presents a mathematical framework for causal nonlinear prediction in settings where observations are generated from an underlying hidden Markov model (HMM). Both the pro…

eess.SY2025

Interacting Particle Systems for Fast Linear Quadratic RL

Anant A Joshi, Heng-Sheng Chang, Amirhossein Taghvaei +2

This paper is concerned with the design of algorithms based on systems of interacting particles to represent, approximate, and learn the optimal control law for reinforcement learn…

cs.LG2025

What can we learn from signals and systems in a transformer? Insights for probabilistic modeling and inference architecture

Heng-Sheng Chang, Prashant G. Mehta

In the 1940s, Wiener introduced a linear predictor, where the future prediction is computed by linearly combining the past data. A transformer generalizes this idea: it is a nonlin…