2 papers
cs.CL2025
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
Di He, Songjun Tu, Ajay Jaiswal +4
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overl…
cs.LG2025
Sebra: Debiasing Through Self-Guided Bias Ranking
Adarsh Kappiyath, Abhra Chaudhuri, Ajay Jaiswal +4
Ranking samples by fine-grained estimates of spuriosity (the degree to which spurious cues are present) has recently been shown to significantly benefit bias mitigation, over the t…