34 citations · 160 across the 32 of their papers we have counts for
5 papers · 1 filter
A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering
Iñigo Alonso, Mirella Lapata
Answering questions over real-world documents requires processing long inputs that interleave text with tables. Optical context compression, which represents context as images, pro…
Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
Adam Fisch, Shubhendu Trivedi, Fantine Huot +5
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist w…
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
Akash Gupta, Amos Storkey, Mirella Lapata
Large Multimodal Models (LMMs) often rely on in-context learning (ICL) to perform new visual question answering (VQA) tasks with minimal supervision. However, ICL performance, espe…
Debating for Better Reasoning: An Unsupervised Multimodal Approach
Ashutosh Adhikari, Mirella Lapata
As Large Language Models (LLMs) gain expertise across diverse domains and modalities, scalable oversight becomes increasingly challenging, particularly when their capabilities may…
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
Louis Mahon, Mirella Lapata
The proliferation of creative video content has driven demand for textual descriptions or summaries that allow users to recall key plot points or get an overview without watching.…