4 papers · 1 filter
Revisiting Gradient Descent: A Dual-Weight Method for Improved Learning
Xi Wang
We introduce a novel framework for learning in neural networks by decomposing each neuron's weight vector into two distinct parts, and , thereby modeling contrastive inf…
Joint control variate for faster black-box variational inference
Xi Wang, Tomas Geffner, Justin Domke
Black-box variational inference performance is sometimes hindered by the use of gradient estimators with high variance. This variance comes from two sources of randomness: Data sub…
Batch size invariant Adam
Xi Wang, Laurence Aitchison
We propose a batch size invariant version of Adam, for use in large-scale, distributed settings, in which the mini-batch is divided into micro-batches which are distributed among w…
Bayesian Low-rank Adaptation for Large Language Models
Adam X. Yang, Maxime Robeyns, Xi Wang +1
Low-rank adaptation (LoRA) has emerged as a new paradigm for cost-efficient fine-tuning of large language models (LLMs). However, fine-tuned LLMs often become overconfident especia…