2 papers
cs.AI2026
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
Qianhong Guo, Wei Xie, Xiaofang Cai +7
Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination,…
cs.IR2024
AutoSurvey: Large Language Models Can Automatically Write Surveys
Yidong Wang, Qi Guo, Wenjin Yao +10
This paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial…