110 citations · 240 across the 26 of their papers we have counts for
24 papers · 1 filter
Buffer-based Gradient Projection for Continual Federated Learning
Shenghong Dai, Jy-yong Sohn, Yicong Chen +5
Continual Federated Learning (CFL) is essential for enabling real-world applications where multiple decentralized clients adaptively learn from continuous data streams. A significa…
Teaching Arithmetic to Small Transformers
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee +2
Large language models like GPT-4 exhibit emergent capabilities across general-purpose tasks, such as basic arithmetic, when trained on extensive text data, even though these tasks…
DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Ying Fan, Olivia Watkins, Yuqing Du +7
Learning from human feedback has been shown to improve text-to-image models. These techniques first learn a reward function that captures what humans care about in the task and the…
Improving Fair Training under Correlation Shifts
Yuji Roh, Kangwook Lee, Steven Euijong Whang +1
Model fairness is an essential element for Trustworthy AI. While many techniques for model fairness have been proposed, most of them assume that the training and deployment data di…
Looped Transformers as Programmable Computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn +3
We present a framework for using transformer networks as universal computers by programming them with specific weights and placing them in a loop. Our input sequence acts as a punc…
Optimizing DDPM Sampling with Shortcut Fine-Tuning
Ying Fan, Kangwook Lee
In this study, we propose Shortcut Fine-Tuning (SFT), a new approach for addressing the challenge of fast sampling of pretrained Denoising Diffusion Probabilistic Models (DDPMs). S…