activity
20242026
collaborators

10 papers

cs.LG2026

Robust Learning of a Group DRO Neuron

Guyang Cao, Shuyao Li, Sushrut Karmalkar +1

We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of c…

stat.ML2026

Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster

Grigory Bartosh, Teodora Pandeva, Sushrut Karmalkar +1

Discrete diffusion models are a powerful class of generative models with strong performance across many domains. For efficiency, however, discrete diffusion typically parameterizes…

cs.LG2026

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task

Alicia Curth, Rachel Lawrence, Sushrut Karmalkar +1

We investigate whether transformers use their depth adaptively across tasks of increasing difficulty. Using a controlled multi-hop relational reasoning task based on family stories…

cs.LG2025

Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing

Iskander Azangulov, Teodora Pandeva, Niranjani Prasad +2

Masked diffusion models (MDMs) offer a compelling alternative to autoregressive models (ARMs) for discrete text generation because they enable parallel token sampling, rather than…

stat.ML2025

A Fourier Space Perspective on Diffusion Models

Fabian Falck, Teodora Pandeva, Kiarash Zahirnia +5

Diffusion models are state-of-the-art generative models on data modalities such as images, audio, proteins and materials. These modalities share the property of exponentially decay…

cs.LG2025

On Learning Parallel Pancakes with Mostly Uniform Weights

Ilias Diakonikolas, Daniel M. Kane, Sushrut Karmalkar +2

We study the complexity of learning -mixtures of Gaussians (-GMMs) on . This task is known to have complexity in full generality. To circumvent this…