2 papers
cs.CL2025
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez +5
Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. (2022) showed that…
cs.CL2025
Towards Zero-Shot Multimodal Machine Translation
Matthieu Futeral, Cordelia Schmid, Benoît Sagot +1
Current multimodal machine translation (MMT) systems rely on fully supervised data (i.e models are trained on sentences with their translations and accompanying images). However, t…