2 citations · 4 across the 10 of their papers we have counts for
Showing 2024 · cs.CVShow all
2 papers · 2 filters
cs.CV2024★ 1 cited
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Gen Luo, Xue Yang, Wenhan Dou +5
In this paper, we focus on monolithic Multimodal Large Language Models (MLLMs) that integrate visual encoding and language decoding into a single LLM. In particular, we identify th…
cs.CV2024
Parameter-Inverted Image Pyramid Networks
Xizhou Zhu, Xue Yang, Zhaokai Wang +6
Image pyramids are commonly used in modern computer vision tasks to obtain multi-scale features for precise understanding of images. However, image pyramids process multiple resolu…