activity
20182024
most citedDocEnTr: An End-to-End Document Image Enhancement Transformer

2 citations · 5 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CV20241 cited

Towards Generative Class Prompt Learning for Fine-grained Visual Recognition

Soumitri Chattopadhyay, Sanket Biswas, Emanuele Vivoli +1

Although foundational vision-language models (VLMs) have proven to be very successful for various semantic discrimination tasks, they still struggle to perform faithfully for fine-…

cs.CV2024

GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding

Nil Biescas, Carlos Boned, Josep Lladós +1

This paper presents GeoContrastNet, a language-agnostic framework to structured document understanding (DU) by integrating a contrastive learning objective with graph attention net…

cs.CV2024

SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition

Adarsh Tiwari, Sanket Biswas, Josep Lladós

We present SketchGPT, a flexible framework that employs a sequence-to-sequence autoregressive model for sketch generation, and completion, and an interpretation case study for sket…

cs.CV20221 cited

A Few Shot Multi-Representation Approach for N-gram Spotting in Historical Manuscripts

Giuseppe De Gregorio, Sanket Biswas, Mohamed Ali Souibgui +4

Despite recent advances in automatic text recognition, the performance remains moderate when it comes to historical manuscripts. This is mainly because of the scarcity of available…

cs.CV20222 cited

DocEnTr: An End-to-End Document Image Enhancement Transformer

Mohamed Ali Souibgui, Sanket Biswas, Sana Khamekhem Jemni +4

Document images can be affected by many degradation scenarios, which cause recognition and processing difficulties. In this age of digitization, it is important to denoise them for…

cs.CV2021

Graph-based Deep Generative Modelling for Document Layout Generation

Sanket Biswas, Pau Riba, Josep Lladós +1

One of the major prerequisites for any deep learning approach is the availability of large-scale training data. When dealing with scanned document images in real world scenarios, t…