1 paper · 1 filter
Haiquan Wang, Yi Chen, Shang Zeng +2
Current evaluations of LLMs in the government domain primarily focus on safety considerations in specific scenarios, while the assessment of the models' own core capabilities, part…