2.5k citations · 2.5k across the 2 of their papers we have counts for
3 papers
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis +5
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larg…
Differential Equation Units: Learning Functional Forms of Activation Functions from Data
MohamadAli Torkamani, Shiv Shankar, Amirmohammad Rooshenas +1
Most deep neural networks use simple, fixed activation functions, such as sigmoids or rectified linear units, regardless of domain or network structure. We introduce differential e…
Learning Compact Neural Networks Using Ordinary Differential Equations as Activation Functions
MohamadAli Torkamani, Phillip Wallis, Shiv Shankar +1
Most deep neural networks use simple, fixed activation functions, such as sigmoids or rectified linear units, regardless of domain or network structure. We introduce differential e…