2 papers
cs.CV2024
Revisiting the Integration of Convolution and Attention for Vision Backbone
Lei Zhu, Xinjiang Wang, Wayne Zhang +1
Convolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate…
cs.CL2024
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
Lei Zhu, Xinjiang Wang, Wayne Zhang +1
A practical large language model (LLM) service may involve a long system prompt, which specifies the instructions, examples, and knowledge documents of the task and is reused acros…