1 citations · 2 across the 12 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Benchmark^2: Systematic Evaluation of LLM Benchmarks
Qi Qian, Chengsong Huang, Jingwen Xu +13
The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose B…
cs.CL2025★ 1 cited
SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
Xinnong Zhang, Jiayu Lin, Xinyi Mou +18
Social simulation is transforming traditional social science research by modeling human behavior through interactions between virtual individuals and their environments. With recen…