5 papers · 1 filter
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
Negar Foroutan, Angelika Romanou, Matin Ansaripour +3
Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic docume…
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
Hernán Maina, Guido Ivetta, Mateo Lione Stuto +3
Visually impaired people could benefit from Visual Question Answering (VQA) systems to interpret text in their surroundings. However, current models often struggle with recognizing…
TANQ: An open domain dataset of table answered questions
Mubashara Akhtar, Chenxi Pang, Andreea Marzoca +2
Language models, potentially augmented with tool usage such as retrieval are becoming the go-to means of answering questions. Understanding and answering questions in real-world se…
Selectively Answering Visual Questions
Julian Martin Eisenschlos, Hernán Maina, Guido Ivetta +1
Recently, large multi-modal models (LMMs) have emerged with the capacity to perform vision tasks such as captioning and visual question answering (VQA) with unprecedented accuracy.…