4 papers
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
Kotaro Yoshida, Atsushi Nitanda
Preconditioning is widely used in machine learning to accelerate convergence on the empirical risk, yet its role on the expected risk remains underexplored. In this work, we invest…
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
Kotaro Yoshida, Yuji Naraki, Takafumi Horie +2
Model merging has emerged as an efficient and flexible paradigm for multi-task learning, with numerous methods being proposed in recent years. However, these state-of-the-art techn…
Robust Invariant Representation Learning by Distribution Extrapolation
Kotaro Yoshida, Konstantinos Slavakis
Invariant risk minimization (IRM) aims to enable out-of-distribution (OOD) generalization in deep learning by learning invariant representations. As IRM poses an inherently challen…
On Fairness of Task Arithmetic: The Role of Task Vectors
Hiroki Naganuma, Kotaro Yoshida, Laura Gomezjurado Gonzalez +3
Model editing techniques, particularly task arithmetic with task vectors, offer an efficient alternative to full fine-tuning by enabling direct parameter updates through simple ari…