Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
Reihaneh Iranmanesh, Saeedeh Davoudi, Pasha Abrishamchian +2
This paper presents a comprehensive evaluation framework for assessing the cultural competence of large language models (LLMs) in Persian. Existing Persian cultural benchmarks rely…
cs.CL2025★ 27 cited
Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE): Manifesto
Christine Bauer, Li Chen, Nicola Ferro +19
During the workshop, we deeply discussed what CONversational Information ACcess (CONIAC) is and its unique features, proposing a world model abstracting it, and defined the Convers…
cs.CL2024
Genetic Approach to Mitigate Hallucination in Generative IR
Hrishikesh Kulkarni, Nazli Goharian, Ophir Frieder +1
Generative language models hallucinate. That is, at times, they generate factually flawed responses. These inaccuracies are particularly insidious because the responses are fluent…