From the 1 of 18 linked papers with an AI index.
4 papers · 1 filter
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
Sushant Mehta
Synthetic data has become essential for training foundation models, yet benchmark contamination threatens evaluation integrity. Although existing detection methods identify token-l…
When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift
Sushant Mehta
Machine learning systems exhibit diverse failure modes: unfairness toward protected groups, brittleness to spurious correlations, poor performance on minority sub-populations, whic…
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
Sushant Mehta, Ishan Gupta
In-context learning (ICL) enables large language models to adapt to new tasks from demonstrations without parameter updates. Despite extensive empirical studies, a principled under…
Muon: Training and Trade-offs with Latent Attention and MoE
Sushant Mehta, Raj Dandekar, Rajat Dandekar +1
We present a comprehensive theoretical and empirical study of the Muon optimizer for training transformers only with a small to medium decoder (30M - 200M parameters), with an emph…