3 papers
cs.CL2026
Multilingual Vision-Language Models, A Survey
Andrei-Alexandru Manea, JindÅich Libovický
This survey examines multilingual vision-language models that process text and images across languages. We review 33 models and 23 benchmarks, spanning encoder-only and generative…
cs.CL2026
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
JindÅich Libovický, JindÅich Helcl, Andrei Manea +1
We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines usin…
cs.CL2025
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
Andrei-Alexandru Manea, JindÅich Libovický
Most pre-trained Vision-Language (VL) models and training data for the downstream tasks are only available in English. Therefore, multilingual VL tasks are solved using cross-lingu…