25 citations · 53 across the 17 of their papers we have counts for
Showing 2021 · cs.CVShow all
2 papers · 2 filters
cs.CV2021
TxT: Crossmodal End-to-End Learning with Transformers
Jan-Martin O. Steitz, Jonas Pfeiffer, Iryna Gurevych +1
Reasoning over multiple modalities, e.g. in Visual Question Answering (VQA), requires an alignment of semantic concepts across domains. Despite the widespread success of end-to-end…
cs.CV2021
Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval
Gregor Geigle, Jonas Pfeiffer, Nils Reimers +2
Current state-of-the-art approaches to cross-modal retrieval process text and visual input jointly, relying on Transformer-based architectures with cross-attention mechanisms that…