5 papers · 1 filter
Muon: Muon with Fractional Spectral Powers
Yihe Dong, Will Sawin
Muon is an increasingly widely used optimizer that replaces a gradient with its polar factor , thereby flattening the singular spectrum. However, full flatten…
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Yihe Dong, Lorenzo Noci, Mikhail Khodak +1
The transformer architecture is central to the success of modern Large Language Models (LLMs), in part due to its surprising ability to perform a wide range of tasks - including ma…
AdaRank: Disagreement Based Module Rank Prediction for Low-rank Adaptation
Yihe Dong
With the rise of language and multimodal models of ever-increasing size, pretraining a general-purpose foundational model and adapting it to downstream tasks has become common prac…
Learned Feature Importance Scores for Automated Feature Engineering
Yihe Dong, Sercan Arik, Nathanael Yoder +1
Feature engineering has demonstrated substantial utility for many machine learning workflows, such as in the small data regime or when distribution shifts are severe. Thus automati…
COSTAR: Improved Temporal Counterfactual Estimation with Self-Supervised Learning
Chuizheng Meng, Yihe Dong, Sercan Ö. Arık +2
Estimation of temporal counterfactual outcomes from observed history is crucial for decision-making in many domains such as healthcare and e-commerce, particularly when randomized…