11 papers
Sample Complexity of Multicalibration for Multilevel Properties
Jiuyao Lu, Krishnakumar Balasubramanian, Aleksandr Podkopaev +1
Calibration requires a predictor to be unbiased after conditioning on its own predictions. Multicalibration asks for this guarantee simultaneously across a collection of groups. Ma…
All Routes Lead to Collapse
K. R. Balasubramanian
Attention sinks, representation collapse, and norm stratification are treated as transformer-specific pathologies. We show they are not specific to attention: they are what content…
FoundCause: Causal Discovery with Latent Confounders from Observational Data
Patrick Blöbaum, Krishnakumar Balasubramanian, Shiva Prasad Kasiviswanathan
Causal discovery from observational data remains challenging due to the need to recover directed structure and latent confounding without interventions. We propose FoundCause, an a…
Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
Krishnakumar Balasubramanian
We propose and analyze a conservative drifting method for one-step generative modeling. The method replaces the original displacement-based drifting velocity by a kernel density es…
Large-Step Training Dynamics of a Two-Factor Linear Transformer Model
Krishnakumar Balasubramanian
Gradient-flow analyses show that simplified linear transformers can learn the in-context linear-regression algorithm, but they do not explain the finite-step behavior of gradient d…
A Quantitative Characterization of Forgetting in Post-Training
Krishnakumar Balasubramanian, Shiva Prasad Kasiviswanathan
Continual post-training of generative models is widely used, yet a principled understanding of when and why forgetting occurs remains limited. We develop theoretical results under…