Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
Jian Chen, Ming Li, Jihyung Kil +6
Most organizational data in this world are stored as documents, and visual retrieval plays a crucial role in unlocking the collective intelligence from all these documents. However…
cs.CV2024
MMR: Evaluating Reading Ability of Large Multimodal Models
Jian Chen, Ruiyi Zhang, Yufan Zhou +3
Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmar…
cs.CV2024
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
Samyadeep Basu, Keivan Rezaei, Priyatham Kattakinda +5
Identifying layers within text-to-image models which control visual attributes can facilitate efficient model editing through closed-form updates. Recent work, leveraging causal tr…