1 citations · 1 across the 3 of their papers we have counts for
3 papers · 1 filter
Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
Raúl Vázquez, Aman Sinha, Chuyuan Li +12
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \te…
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
Yihao Ding, Lorenzo Vaiani, Caren Han +4
This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is de…
Enhancing BERT-Based Visual Question Answering through Keyword-Driven Sentence Selection
Davide Napolitano, Lorenzo Vaiani, Luca Cagliero
The Document-based Visual Question Answering competition addresses the automatic detection of parent-child relationships between elements in multi-page documents. The goal is to id…