4 papers
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
Shahar Levy, Eliya Habba, Reshef Mintz +3
Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually de…
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
Gili Lior, Eliya Habba, Shahar Levy +2
LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations…
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
Shahar Levy, Nir Mazor, Lihi Shalmon +2
Retrieval-Augmented Generation (RAG) enhances the accuracy of Large Language Model (LLM) responses by leveraging relevant external documents during generation. Although previous st…
SEAM: A Stochastic Benchmark for Multi-Document Tasks
Gili Lior, Avi Caciularu, Arie Cattan +3
Various tasks, such as summarization, multi-hop question answering, or coreference resolution, are naturally phrased over collections of real-world documents. Such tasks present a…