Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Chang Ma, Junlei Zhang, Zhihao Zhu +6
Evaluating Large Language Models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications.…
cs.CL2024
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology
Junlei Zhang, Hongliang He, Nirui Song +8
The critical field of psychology necessitates a comprehensive benchmark to enhance the evaluation and development of domain-specific Large Language Models (LLMs). Existing MMLU-typ…