activity
20192026
most citedBandwidth Enables Generalization in Quantum Kernel Models

19 citations · 120 across the 55 of their papers we have counts for

collaborators
Showing cs.LGShow all

30 papers · 1 filter

cs.LG2026

On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

William L. Tong, Aryo Lotfi, Emmanuel Abbe +6

State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reason…

cs.LG2026

A Defense of the Quadratic Model

Alexandru Meterez, Pranav Ajit Nair, Depen Morwani +3

Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically trac…

cs.LG2026

Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover

Indranil Halder, Annesya Banerjee, Cengiz Pehlevan

Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. Empirically, we find that adversarial prompt-injection attacks can amplify attac…

cs.LG2026

Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging

Alexandru Meterez, Pranav Ajit Nair, Depen Morwani +2

Large language models are increasingly trained in continual or open-ended settings, where the total training horizon is not known in advance. Despite this, most existing pretrainin…

cs.LG2026

Universal One-third Time Scaling in Learning Peaked Distributions

Yizhou Liu, Ziming Liu, Cengiz Pehlevan +1

Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic a…

cs.LG2026

A Random Matrix Theory Perspective on the Consistency of Diffusion Models

Binxu Wang, Jacob Zavatone-Veth, Cengiz Pehlevan

Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed. We trace this consistency to a…