5 papers · 1 filter
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn +3
Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predomi…
The PLLuM Instruction Corpus
Piotr Pęzik, Filip Żarnecki, Konrad Kaczyński +50
This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…
Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse
Anna Kołos, Katarzyna Lorenc, Emilia Wiśnios +1
The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. W…
StyloMetrix: An Open-Source Multilingual Tool for Representing Stylometric Vectors
Inez Okulska, Daria Stetsenko, Anna Kołos +3
This work aims to provide an overview on the open-source multilanguage tool called StyloMetrix. It offers stylometric text representations that cover various aspects of grammar, sy…
BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service
Anna Kołos, Inez Okulska, Kinga Głąbińska +4
Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require impr…