4 papers
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
Ivan KartáÄ, Kristýna Onderková, Jan Bronec +3
This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an efficient modular neuro-symbol…
Reasoning Gets Harder for LLMs Inside A Dialogue
Ivan KartáÄ, Mateusz Lango, OndÅej DuÅ¡ek
Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in t…
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
ZdenÄk Kasner, Vilém Zouhar, PatrÃcia Schmidtová +7
Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
Ivan KartáÄ, Mateusz Lango, OndÅej DuÅ¡ek
Large Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, exist…