12 citations · 17 across the 4 of their papers we have counts for
4 papers
Vcc: Scaling Transformers to 128K Tokens or More by Prioritizing Important Tokens
Zhanpeng Zeng, Cole Hawkins, Mingyi Hong +4
Transformers are central in modern natural language processing and computer vision applications. Despite recent works devoted to reducing the quadratic cost of such models (as a fu…
Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition
Shuhuai Ren, Aston Zhang, Yi Zhu +5
This work proposes POMP, a prompt pre-training method for vision-language models. Being memory and computation efficient, POMP enables the learned prompt to condense semantic infor…
SPT: Semi-Parametric Prompt Tuning for Multitask Prompted Learning
M Saiful Bari, Aston Zhang, Shuai Zheng +4
Pre-trained large language models can efficiently interpolate human-written prompts in a natural way. Multitask prompted learning can help generalization through a diverse set of t…
SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing
Chaoyang He, Shuai Zheng, Aston Zhang +4
The mixture of Expert (MoE) parallelism is a recent advancement that scales up the model size with constant computational cost. MoE selects different sets of parameters (i.e., expe…