2 papers
cs.IR2026
EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge
Yuhong Sun, Joachim Rahmfeld, Chris Weaver +4
Retrieval-Augmented Generation (RAG) has become the standard approach for grounding large language models in information that was not available during training. While existing data…
cs.CL2025
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
Yuhong Sun, Zhangyue Yin, Xuanjing Huang +2
Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains. Math Word Problems (MWPs) serve as a crucial benchmark for evaluating LLMs' reasoning…