40 citations · 40 across the 2 of their papers we have counts for
2 papers
cs.CL2023
LAIT: Efficient Multi-Segment Encoding in Transformers with Layer-Adjustable Interaction
Jeremiah Milbauer, Annie Louis, Mohammad Javad Hosseini +3
Transformer encoders contextualize token representations by attending to all other tokens at each layer, leading to quadratic increase in compute effort with the input length. In p…
cs.CL2022★ 40 cited
Confident Adaptive Language Modeling
Tal Schuster, Adam Fisch, Jai Gupta +5
Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks. These gains come with a drastic increase in th…