evaluation metrics 1evidence reliability 1multi-step question answering 1search agents 1stress testing 1synthetic retrieval 1
From the 1 of 7 linked papers with an AI index.
Showing 2025Show all
2 papers · 1 filter
cs.LG2025
Statistical Deficiency for Task Inclusion Estimation
Loïc Fosse, Frédéric Béchet, Benoît Favre +5
Tasks are central in machine learning, as they are the most natural objects to assess the capabilities of current models. The trend is to build general models able to address any t…
cs.CL2025
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant +2
This paper explores the robustness of language models (LMs) to variations in the temporal context within factual knowledge. It examines whether LMs can correctly associate a tempor…