Publications (15)
CLIPTER: Looking at the Bigger Picture in Scene Text Recognition
Aviad Aberdam, David Bensaïd, Alona Golts +5
Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizer…
Out-of-Vocabulary Challenge Report
Sergi Garcia-Bordils, Andrés Mafla, Ali Furkan Biten +5
This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Re…
ScrabbleGAN: Semi-Supervised Varying Length Handwritten Text Generation
Sharon Fogel, Hadar Averbuch-Elor, Sarel Cohen +2
Optical character recognition (OCR) systems performance have improved significantly in the deep learning era. This is especially true for handwritten text recognition (HTR), where…
Question Aware Vision Transformer for Multimodal Reasoning
Roy Ganz, Yair Kittenplon, Aviad Aberdam +4
Vision-Language (VL) models have gained significant research focus, enabling remarkable advances in multimodal reasoning. These architectures typically comprise a vision encoder, a…
Towards Models that Can See and Read
Roy Ganz, Oren Nuriel, Aviad Aberdam +3
Visual Question Answering (VQA) and Image Captioning (CAP), which are among the most popular vision-language tasks, have analogous scene-text versions that require reasoning from t…
DODO: Discrete OCR Diffusion Models
Sean Man, Gilad Deutch, Roy Ganz +4
Optical Character Recognition (OCR) is a fundamental task for digitizing information, serving as a critical bridge between visual data and textual understanding. While modern Visio…
Learning Multimodal Affinities for Textual Editing in Images
Or Perel, Oron Anschel, Omri Ben-Eliezer +2
Nowadays, as cameras are rapidly adopted in our daily routine, images of documents are becoming both abundant and prevalent. Unlike natural images that capture physical objects, do…
VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
Ofir Abramovich, Niv Nayman, Sharon Fogel +7
In recent years, notable advancements have been made in the domain of visual document understanding, with the prevailing architecture comprising a cascade of vision and language mo…
DocVLM: Make Your VLM an Efficient Reader
Mor Shpigel Nacson, Aviad Aberdam, Roy Ganz +5
Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks p…
On Calibration of Scene-Text Recognition Models
Ron Slossberg, Oron Anschel, Amir Markovitz +6
In this work, we study the problem of word-level confidence calibration for scene-text recognition (STR). Although the topic of confidence calibration has been an active research a…
Bayesian Time-of-Flight for Realtime Shape, Illumination and Albedo
Amit Adam, Christoph Dann, Omer Yair +2
We propose a computational model for shape, illumination and albedo inference in a pulsed time-of-flight (TOF) camera. In contrast to TOF cameras based on phase modulation, our cam…
SCATTER: Selective Context Attentional Scene Text Recognizer
Ron Litman, Oron Anschel, Shahar Tsiper +3
Scene Text Recognition (STR), the task of recognizing text against complex image backgrounds, is an active area of research. Current state-of-the-art (SOTA) methods still struggle…
Can You Read Me Now? Content Aware Rectification using Angle Supervision
Amir Markovitz, Inbal Lavi, Or Perel +2
The ubiquity of smartphone cameras has led to more and more documents being captured by cameras rather than scanned. Unlike flatbed scanners, photographed documents are often folde…
Multimodal Semi-Supervised Learning for Text Recognition
Aviad Aberdam, Roy Ganz, Shai Mazor +1
Until recently, the number of public real-world text images was insufficient for training scene text recognizers. Therefore, most modern training methods rely on synthetic data and…
Sequence-to-Sequence Contrastive Learning for Text Recognition
Aviad Aberdam, Ron Litman, Shahar Tsiper +5
We propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence…