Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law
Ethan Zhao, Maksym Taranukhin, Wei Cui +2
RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmarks have been developed to ev…
cs.CL2025
FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
Bowen Qin, Chen Yue, Fang Yin +26
We conduct a moderate-scale contamination-free (to some extent) evaluation of current large reasoning models (LRMs) with some preliminary findings. We also release ROME, our evalua…