3 papers
cs.LG2026
Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators
Anupama Sridhar, Alexander Johansen
Long chain-of-thought reasoning and agentic tool-calling produce traces spanning tens of thousands of tokens, yet Transformer KV caches grow linearly with sequence length, creating…
stat.ML2025
Convergence of Adam in Deep ReLU Networks via Directional Complexity and Kakeya Bounds
Anupama Sridhar, Alexander Johansen
First-order adaptive optimization methods like Adam are the default choices for training modern deep neural networks. Despite their empirical success, the theoretical understanding…
stat.ML2025
Convergence of TD(0) under Polynomial Mixing with Nonlinear Function Approximation
Anupama Sridhar, Alexander Johansen
Temporal Difference Learning (TD(0)) is fundamental in reinforcement learning, yet its finite-sample behavior under non-i.i.d. data and nonlinear approximation remains unknown. We…