Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
Junru Song, Yimeng Hu, Yijing Chen +4
Large Vision-Language Models (VLMs) have achieved remarkable multimodal performance yet remain prone to factual hallucinations, particularly in long-tail or specialized domains. Mo…
cs.CV2025
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
Tianfan Peng, Yuntao Du, Pengzhou Ji +9
Large multimodal models (LMMs) often suffer from severe inference inefficiency due to the large number of visual tokens introduced by image encoders. While recent token compression…