most citedWhat is the best model? Application-driven Evaluation for Large Language Models

1 citations · 2 across the 2 of their papers we have counts for

collaborators

6 papers

cs.CL2025

Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts

Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu +11

DeepSeek-R1, renowned for its exceptional reasoning capabilities and open-source strategy, is significantly influencing the global artificial intelligence landscape. However, it ex…

cs.AI2025

Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis

Kaikai Zhao, Zhaoxiang Liu, Xuejiao Lei +12

DeepSeek-R1, known for its low training cost and exceptional reasoning capabilities, has achieved state-of-the-art performance on various benchmarks. However, detailed evaluations…

cs.CL2025

Safety Evaluation of DeepSeek Models in Chinese Contexts

Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu +8

Recently, the DeepSeek series of models, leveraging their exceptional reasoning capabilities and open-source strategy, is reshaping the global AI landscape. Despite these advantage…

cs.CL20241 cited

Methodology of Adapting Large English Language Models for Specific Cultural Contexts

Wenjing Zhang, Siqi Xiao, Xuejiao Lei +7

The rapid growth of large language models(LLMs) has emerged as a prominent trend in the field of artificial intelligence. However, current state-of-the-art LLMs are predominantly b…

cs.CL20241 cited

What is the best model? Application-driven Evaluation for Large Language Models

Shiguo Lian, Kaikai Zhao, Xinhui Liu +5

General large language models enhanced with supervised fine-tuning and reinforcement learning from human feedback are increasingly popular in academia and industry as they generali…

cs.CL2024

CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu +5

With the profound development of large language models(LLMs), their safety concerns have garnered increasing attention. However, there is a scarcity of Chinese safety benchmarks fo…