6 papers
Shieldstral
Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +273
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…
Can Models Help Us Create Better Models? Evaluating LLMs as Data Scientists
Michał Pietruszka, Łukasz Borchmann, Aleksander Jędrosz +1
We present a benchmark for large language models designed to tackle one of the most knowledge-intensive tasks in data science: writing feature engineering code, which requires doma…
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz +3
We address the challenging problem of Natural Language Comprehension beyond plain-text documents by introducing the TILT neural network architecture which simultaneously learns lay…
From Dataset Recycling to Multi-Property Extraction and Beyond
Tomasz Dwojak, Michał Pietruszka, Łukasz Borchmann +2
This paper investigates various Transformer architectures on the WikiReading Information Extraction and Machine Reading Comprehension dataset. The proposed dual-source model outper…
Successive Halving Top-k Operator
Michał Pietruszka, Łukasz Borchmann, Filip Graliński
We propose a differentiable successive halving method of relaxing the top-k operator, rendering gradient-based optimization possible. The need to perform softmax iteratively on the…
On the Multi-Property Extraction and Beyond
Tomasz Dwojak, Michał Pietruszka, Łukasz Borchmann +2
In this paper, we investigate the Dual-source Transformer architecture on the WikiReading information extraction and machine reading comprehension dataset. The proposed model outpe…