2 papers
cs.LG2025
Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms
Ruoxin Zhang, Zhizhao Wen, Chao Wang +3
With the rapid evolution of large language models, retrieval enhanced generation technology has been widely used due to its ability to integrate external knowledge to improve outpu…
cs.CL2025
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Shi Qiu, Shaoyang Guo, Zhuo-Yang Song +51
Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed e…