2 papers
cs.CL2026
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
Yuchen Lu, Run Yang, Yichen Zhang +6
Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-dri…
physics.soc-ph2025
Generalized Multi-agent Social Simulation Framework
Gang Li, Jie Lin, Yining Tang +6
Multi-agent social interaction has clearly benefited from Large Language Models. However, current simulation systems still face challenges such as difficulties in scaling to divers…