activity
20242026
collaborators

5 papers

cs.IR2026

UnWeaving the knots of GraphRAG -- turns out VectorRAG is almost enough

Ryszard Tuora, Mateusz Galiński, Michał Godziszewski +4

One of the key problems in Retrieval-augmented generation (RAG) systems is that chunk-based retrieval pipelines represent the source chunks as atomic objects, mixing the informatio…

cs.AI2026

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering

Mateusz Czyżnikiewicz, Ryszard Tuora, Adam Kozakiewicz +6

Retrieval-Augmented Generation (RAG) systems for question answering typically retrieve evidence by semantic similarity between the query and document chunks. While effective for un…

cs.SD2026

A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models

Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski +4

The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identific…

cs.CL2025

Preservation of Language Understanding Capabilities in Speech-aware Large Language Models

Marek Kubis, Paweł Skórzewski, Iwona Christop +4

The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes tex…

eess.AS2024

Augmenting Polish Automatic Speech Recognition System With Synthetic Data

Łukasz Bondaruk, Jakub Kubiak, Mateusz Czyżnikiewicz

This paper presents a system developed for submission to Poleval 2024, Task 3: Polish Automatic Speech Recognition Challenge. We describe Voicebox-based speech synthesis pipeline a…