2 papers
cs.LG2025
Implicit Regularization Makes Overparameterized Asymmetric Matrix Sensing Robust to Perturbations
Johan S. Wind
Several key questions remain unanswered regarding overparameterized learning models. It is unclear how (stochastic) gradient descent finds solutions that generalize well, and in pa…
cs.CL2025
RWKV-7 "Goose" with Expressive Dynamic State Evolution
Bo Peng, Ruichong Zhang, Daniel Goldstein +15
We present RWKV-7 "Goose", a new sequence modeling architecture with constant memory usage and constant inference time per token. Despite being trained on dramatically fewer tokens…