collaborators

12 papers

stat.ML2026

The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex Optimization

Nived Rajaraman, Yanjun Han

We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for -Lipschitz functions on the -dimensional Euclidean ball. For ti…

cs.LG2026

Learning to Reason with Curriculum II: Compositional Generalization

Nived Rajaraman, Audrey Huang, Miroslav Dudik +3

Compositional generalization, the ability to solve complex problems by combining solutions to simpler sub-problems, is a fundamental capability of both natural and artificial intel…

cs.LG2026

From Markov to Laplace: How Mamba In-Context Learns Markov Chains

Marco Bondaschi, Nived Rajaraman, Xiuying Wei +5

While transformer-based language models have driven the AI revolution thus far, their computational complexity has spurred growing interest in viable alternatives, such as structur…

cs.LG2026

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning

Akshay Krishnamurthy, Audrey Huang, Nived Rajaraman

Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We st…

cs.LG2026

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

Nived Rajaraman, Audrey Huang, Miro Dudik +3

Chain-of-thought reasoning, where language models expend additional computation by producing thinking tokens prior to final responses, has driven significant advances in model capa…

stat.ML2026

Interactive Learning of Single-Index Models via Stochastic Gradient Descent

Nived Rajaraman, Yanjun Han

Stochastic gradient descent (SGD) is a cornerstone algorithm for high-dimensional optimization, renowned for its empirical successes. Recent theoretical advances have provided a de…