4 citations · 4 across the 2 of their papers we have counts for
3 papers
CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data
Michał Turski, Tomasz Stanisławek, Karol Kaczmarek +2
In recent years, the field of document understanding has progressed a lot. A significant part of this progress has been possible thanks to the use of language models pretrained on…
STable: Table Generation Framework for Encoder-Decoder Models
Michał Pietruszka, Michał Turski, Łukasz Borchmann +5
The output structure of database-like tables, consisting of values structured in horizontal rows and vertical columns identifiable by name, can cover a wide range of NLP tasks. Fol…
LAMBERT: Layout-Aware (Language) Modeling for information extraction
Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek +4
We introduce a simple new approach to the problem of understanding documents where non-trivial layout influences the local semantics. To this end, we modify the Transformer encoder…