1 paper · 1 filter
John Gkountouras, Ivan Titov
Recent text-only models demonstrate remarkable mathematical reasoning capabilities. Extending these to visual domains requires vision-language models to translate images into text…