collaborators

5 papers

cs.CL2026

VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations

Ivan Kartáč, Ivan Kartáč, Jan Tovarys +3

Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as…

cs.CL2026

UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning

Ivan Kartáč, Kristýna Onderková, Jan Bronec +3

This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an efficient modular neuro-symbol…

cs.CL2026

Reasoning Gets Harder for LLMs Inside A Dialogue

Ivan Kartáč, Mateusz Lango, Ondřej Dušek

Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in t…

cs.CL2026

LLMs as Span Annotators: A Comparative Study of LLMs and Humans

Zdeněk Kasner, Vilém Zouhar, Patrícia Schmidtová +7

Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…

cs.CL2025

OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs

Ivan Kartáč, Mateusz Lango, Ondřej Dušek

Large Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, exist…