1 paper · 1 filter
Hao Zhang, Mengsi Lyu, Bo Huang +2
Large Multimodal Models (LMMs) have proven effective on various tasks. They typically encode visual inputs into Original Model sequences of tokens, which are then concatenated with…