12 papers
VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations
Ivan KartáÄ, Ivan Kartáč, Jan Tovarys +3
Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as…
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resourc…
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
Ivan KartáÄ, Kristýna Onderková, Jan Bronec +3
This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an efficient modular neuro-symbol…
Reasoning Gets Harder for LLMs Inside A Dialogue
Ivan KartáÄ, Mateusz Lango, OndÅej DuÅ¡ek
Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in t…
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
Wiktor Kamzela, Mateusz Lango, Ondrej Dusek
In this paper, we use large language models to generate personalized stories for language learners, using only the vocabulary they know. The generated texts are specifically writte…
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
Mateusz Lango, OndÅej DuÅ¡ek
We present a novel neurosymbolic framework for RDF-to-text generation, in which the model is "trained" through collaborative interactions among multiple LLM agents rather than trad…