13 citations · 59 across the 33 of their papers we have counts for
1 paper · 1 filter
Shaolei Zhang, Qingkai Fang, Zhe Yang +1
The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision to…