5 papers
T^2MLR: Transformer with Temporal Middle-Layer Recurrence
Ziyang Cai, Xingyu Zhu, Yihe Dong +2
Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reason…
Muon: Muon with Fractional Spectral Powers
Yihe Dong, Will Sawin
Muon is an increasingly widely used optimizer that replaces a gradient with its polar factor , thereby flattening the singular spectrum. However, full flatten…
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Yihe Dong, Lorenzo Noci, Mikhail Khodak +1
The transformer architecture is central to the success of modern Large Language Models (LLMs), in part due to its surprising ability to perform a wide range of tasks - including ma…
AdaRank: Disagreement Based Module Rank Prediction for Low-rank Adaptation
Yihe Dong
With the rise of language and multimodal models of ever-increasing size, pretraining a general-purpose foundational model and adapting it to downstream tasks has become common prac…
Learned Feature Importance Scores for Automated Feature Engineering
Yihe Dong, Sercan Arik, Nathanael Yoder +1
Feature engineering has demonstrated substantial utility for many machine learning workflows, such as in the small data regime or when distribution shifts are severe. Thus automati…