From the 1 of 136 linked papers with an AI index.
54 citations · 56 across the 37 of their papers we have counts for
3 papers · 1 filter
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
Haonan Chen, Sicheng Gao, Radu Timofte +2
Modern information systems often involve different types of items, e.g., a text query, an image, a video clip, or an audio segment. This motivates omni-modal embedding models that…
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
Mohamed Gado, Towhid Taliee, Muhammad Memon +2
Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper pre…
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
Gregor Geigle, Florian Schneider, Carolin Holtermann +4
Most Large Vision-Language Models (LVLMs) to date are trained predominantly on English data, which makes them struggle to understand non-English input and fail to generate output i…