2 papers
cs.CV2025
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
Tianfan Peng, Yuntao Du, Pengzhou Ji +9
Large multimodal models (LMMs) often suffer from severe inference inefficiency due to the large number of visual tokens introduced by image encoders. While recent token compression…
cs.CL2025
MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models
Kailin Jiang, Ning Jiang, Yuntao Du +8
Large Multimodal Models (LMMs) encode rich factual knowledge via cross-modal pre-training, yet their static representations struggle to maintain an accurate understanding of time-s…