3 papers
cs.LG2025
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
Carlos Boned Riera, David Romero Sanchez, Oriol Ramos Terrades
In recent years, increasingly large models have achieved outstanding performance across CV tasks. However, these models demand substantial computational resources and storage, and…
cs.CV2024
Recurrent Few-Shot model for Document Verification
Maxime Talarmain, Carlos Boned, Sanket Biswas +1
General-purpose ID, or travel, document image- and video-based verification systems have yet to achieve good enough performance to be considered a solved problem. There are several…
cs.CV2024
GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
Nil Biescas, Carlos Boned, Josep Lladós +1
This paper presents GeoContrastNet, a language-agnostic framework to structured document understanding (DU) by integrating a contrastive learning objective with graph attention net…