20 citations · 21 across the 5 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
Mikail Khona, Aditya Vavre, Boxiang Wang +11
Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. I…
cs.LG2026
TorchKM: A GPU-Oriented Library for Kernel Learning and Model Selection
Yikai Zhang, Gaoxiang Jia, Jie Ding +1
TorchKM is an open-source library for kernel machines, including support vector machines, kernel logistic regression, and kernel quantile regression, with GPU acceleration. The lib…
cs.LG2021
Partially Interpretable Estimators (PIE): Black-Box-Refined Interpretable Machine Learning
Tong Wang, Jingyi Yang, Yunyi Li +1
We propose Partially Interpretable Estimators (PIE) which attribute a prediction to individual features via an interpretable model, while a (possibly) small part of the PIE predict…