2 papers
cs.CL2026
How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation
Elitsa Yotkova, Violeta Kastreva, Petar Velkov +4
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be co…
cs.CL2026
Schützen: Evaluating LLM Safety in Bulgarian and German Contexts
Kiril Georgiev, Yuxia Wang, Dimitar Iliyanov Dimitrov +2
Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disrespectful content. Although…