2 papers
cs.CV2025
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
Vittorio Pippi, Matthieu Guillaumin, Silvia Cascianelli +3
Large Multimodal Models (LMMs) are powerful tools that are capable of reasoning and understanding multimodal information beyond text and language. Despite their entrenched impact,…
cs.CV2025
UniCoRN: Unified Commented Retrieval Network with LMMs
Maximilian Jaritz, Matthieu Guillaumin, Sabine Sternig +1
Multimodal retrieval methods have limitations in handling complex, compositional queries that require reasoning about the visual content of both the query and the retrieved entitie…