Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Law of Vision Representation in MLLMs
Shijia Yang, Bohan Zhai, Quanzeng You +3
We present the "Law of Vision Representation" in multimodal large language models (MLLMs). It reveals a strong correlation between the combination of cross-modal alignment, corresp…
cs.CV2024
Learning Stackable and Skippable LEGO Bricks for Efficient, Reconfigurable, and Variable-Resolution Diffusion Modeling
Huangjie Zheng, Zhendong Wang, Jianbo Yuan +5
Diffusion models excel at generating photo-realistic images but come with significant computational costs in both training and sampling. While various techniques address these comp…