1 paper
Dingchen Yang, Bowen Cao, Anran Zhang +3
Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational b…