activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn +3

Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predomi…

cs.CL2025

The PLLuM Instruction Corpus

Piotr Pęzik, Filip Żarnecki, Konrad Kaczyński +50

This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…

cs.CL2024

Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse

Anna Kołos, Katarzyna Lorenc, Emilia Wiśnios +1

The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. W…

cs.CL20231 cited

StyloMetrix: An Open-Source Multilingual Tool for Representing Stylometric Vectors

Inez Okulska, Daria Stetsenko, Anna Kołos +3

This work aims to provide an overview on the open-source multilanguage tool called StyloMetrix. It offers stylometric text representations that cover various aspects of grammar, sy…

cs.CL2023

BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service

Anna Kołos, Inez Okulska, Kinga Głąbińska +4

Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require impr…