3 papers
cs.AI2026
A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering
Iñigo Alonso, Mirella Lapata
Answering questions over real-world documents requires processing long inputs that interleave text with tables. Optical context compression, which represents context as images, pro…
cs.CV2025
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
Iñigo Alonso, Imanol Miranda, Eneko Agirre +1
While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of rea…
cs.CL2025
Vision-Language Models Struggle to Align Entities across Modalities
Iñigo Alonso, Gorka Azkune, Ander Salaberria +2
Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed…