7 papers
AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis
Saeedeh Davoudi, Reihaneh Iranmanesh, Ophir Frieder +1
Medical knowledge is continuously evolving. This creates a need to update or selectively forget information encoded in already-trained medical LLMs. Machine unlearning aims to remo…
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
Reihaneh Iranmanesh, Saeedeh Davoudi, Pasha Abrishamchian +2
This paper presents a comprehensive evaluation framework for assessing the cultural competence of large language models (LLMs) in Persian. Existing Persian cultural benchmarks rely…
A Picture of Agentic Search
Francesca Pezzuti, Ophir Frieder, Fabrizio Silvestri +2
With automated systems increasingly issuing search queries alongside humans, Information Retrieval (IR) faces a major shift. Yet IR remains human-centred, with systems, evaluation…
DRAMA: Domain Retrieval using Adaptive Module Allocation
Pranav Kasela, Marco Braga, Ophir Frieder +3
Neural models are increasingly used in Web-scale Information Retrieval (IR). However, relying on these models introduces substantial computational and energy requirements, leading…
Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE): Manifesto
Christine Bauer, Li Chen, Nicola Ferro +19
During the workshop, we deeply discussed what CONversational Information ACcess (CONIAC) is and its unique features, proposing a world model abstracting it, and defined the Convers…
Intercept Cancer: Cancer Pre-Screening with Large Scale Healthcare Foundation Models
Liwen Sun, Hao-Ren Yao, Gary Gao +2
Cancer screening, leading to early detection, saves lives. Unfortunately, existing screening techniques require expensive and intrusive medical procedures, not globally available,…