5 papers
Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering
Xianliang Li, Zihan Zhang, Weiyang Liu +1
Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear. Existing analyses of Muo…
Many-to-Many Matching via Sparsity Controlled Optimal Transport
Weijie Liu, Han Bao, Makoto Yamada +3
Many-to-many matching seeks to match multiple points in one set and multiple points in another set, which is a basis for a wide range of data mining problems. It can be naturally r…
PhiNets: Brain-inspired Non-contrastive Learning Based on Temporal Prediction Hypothesis
Satoki Ishikawa, Makoto Yamada, Han Bao +1
Predictive coding is a theory which hypothesises that cortex predicts sensory inputs at various levels of abstraction to minimise prediction errors. Inspired by predictive coding,…
Necessary and Sufficient Watermark for Large Language Models
Yuki Takezawa, Ryoma Sato, Han Bao +2
In recent years, large language models (LLMs) have achieved remarkable performances in various NLP tasks. They can generate texts that are indistinguishable from those written by h…
Parameter-free Clipped Gradient Descent Meets Polyak
Yuki Takezawa, Han Bao, Ryoma Sato +2
Gradient descent and its variants are de facto standard algorithms for training machine learning models. As gradient descent is sensitive to its hyperparameters, we need to tune th…