11 citations · 14 across the 3 of their papers we have counts for
3 papers
cs.CV2024
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
Lingchen Meng, Jianwei Yang, Rui Tian +4
Most large multimodal models (LMMs) are implemented by feeding visual tokens as a sequence into the first layer of a large language model (LLM). The resulting architecture is simpl…
cs.CV2024★ 11 cited
Rewrite the Stars
Xu Ma, Xiyang Dai, Yue Bai +2
Recent studies have drawn attention to the untapped potential of the "star operation" (element-wise multiplication) in network design. While intuitive explanations abound, the foun…
cs.CV2024★ 3 cited
Efficient Modulation for Vision Networks
Xu Ma, Xiyang Dai, Jianwei Yang +4
In this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context…