2 papers
cs.AI2026
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
Sihang Jiang, Lipeng Ma, Zhonghua Hong +9
Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience…
cs.CL2025
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
Jinghao Zhang, Sihang Jiang, Shiwei Guo +7
As large language models (LLMs) are increasingly deployed in diverse cultural environments, evaluating their cultural understanding capability has become essential for ensuring tru…