Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
Xianjin Wu, Dingkang Liang, Tianrui Feng +5
While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and…
cs.CV2024
EVLM: An Efficient Vision-Language Model for Visual Understanding
Kaibing Chen, Dong Shen, Hanwen Zhong +14
In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…