3 papers
cs.LG2026
Faster Query-Key Learning Sharpens Attention in Self-Attention Models
Rahul Vashisht, Harish G. Ramaswamy
A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended repre…
cs.LG2025
Optimizing Shortfall Risk Metric for Learning Regression Models
Harish G. Ramaswamy, L. A. Prashanth
We consider the problem of estimating and optimizing utility-based shortfall risk (UBSR) of a loss, say , in the context of a regression problem. Empirical risk min…
cs.LG2024
Impact of Label Noise on Learning Complex Features
Rahul Vashisht, P. Krishna Kumar, Harsha Vardhan Govind +1
Neural networks trained with stochastic gradient descent exhibit an inductive bias towards simpler decision boundaries, typically converging to a narrow family of functions, and of…