4 papers
How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation
Chase M. Fensore, Kaustubh Dhole, Jason Fan +2
Evaluating retrieval-augmented generation (RAG) systems requires benchmarks that capture diverse question characteristics, yet practitioners lack empirical guidance on which dimens…
Anatomy of a Query: W5H Dimensions and FAR Patterns for Text-to-SQL Evaluation
Vicki Stover Hertzberg, Eduardo Valverde, Joyce C. Ho
Natural language interfaces to databases have gained popularity, yet the theoretical foundations for evaluating and designing these systems remain underdeveloped. We present QUEST…
Reasoning Traces Shape Outputs but Models Won't Say So
Yijie Hao, Lingjie Chen, Ali Emami +1
Can we trust the reasoning traces that large reasoning models (LRMs) produce? We investigate whether these traces faithfully reflect what drives model outputs, and whether models w…
Evaluating Hybrid Retrieval Augmented Generation using Dynamic Test Sets: LiveRAG Challenge
Chase Fensore, Kaustubh Dhole, Joyce C Ho +1
We present our submission to the LiveRAG Challenge 2025, which evaluates retrieval-augmented generation (RAG) systems on dynamic test sets using the FineWeb-10BT corpus. Our final…