3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2024
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
Lingchen Meng, Jianwei Yang, Rui Tian +4
Most large multimodal models (LMMs) are implemented by feeding visual tokens as a sequence into the first layer of a large language model (LLM). The resulting architecture is simpl…
cs.CV2024★ 3 cited
Efficient Modulation for Vision Networks
Xu Ma, Xiyang Dai, Jianwei Yang +4
In this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context…