5 papers
Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial Participation
Kaoru Otsuka, Yuki Takezawa, Makoto Yamada
Partial participation is essential for communication-efficient federated learning at scale, yet existing Byzantine-robust methods typically assume full client participation. In the…
Any-stepsize Gradient Descent for Separable Data under Fenchel-Young Losses
Han Bao, Shinsaku Sakaue, Yuki Takezawa
The gradient descent (GD) has been one of the most common optimizer in machine learning. In particular, the loss landscape of a neural network is typically sharpened during the ini…
PhiNets: Brain-inspired Non-contrastive Learning Based on Temporal Prediction Hypothesis
Satoki Ishikawa, Makoto Yamada, Han Bao +1
Predictive coding is a theory which hypothesises that cortex predicts sensory inputs at various levels of abstraction to minimise prediction errors. Inspired by predictive coding,…
Necessary and Sufficient Watermark for Large Language Models
Yuki Takezawa, Ryoma Sato, Han Bao +2
In recent years, large language models (LLMs) have achieved remarkable performances in various NLP tasks. They can generate texts that are indistinguishable from those written by h…
Parameter-free Clipped Gradient Descent Meets Polyak
Yuki Takezawa, Han Bao, Ryoma Sato +2
Gradient descent and its variants are de facto standard algorithms for training machine learning models. As gradient descent is sensitive to its hyperparameters, we need to tune th…