most citedSafety Assessment of Chinese Large Language Models

16 citations · 44 across the 14 of their papers we have counts for

collaborators

23 papers

cs.CL20241 cited

LegalAgentBench: Evaluating LLM Agents in Legal Domain

Haitao Li, Junjie Chen, Jingli Yang +10

With the increasing intelligence and autonomy of LLM agents, their potential applications in the legal domain are becoming increasingly apparent. However, existing general-domain b…

cs.CL2024

The Superalignment of Superhuman Intelligence with Large Language Models

Minlie Huang, Yingkang Wang, Shiyao Cui +2

We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becom…

cs.CL2024

CharacterBench: Benchmarking Character Customization of Large Language Models

Jinfeng Zhou, Yongkang Huang, Bosi Wen +13

Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character c…

cs.CL2024

Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework

Xuanming Zhang, Yuxuan Chen, Yiming Zheng +3

In real world software development, improper or missing exception handling can severely impact the robustness and reliability of code. Exception handling mechanisms require develop…

cs.CL2024

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Zhenyu Hou, Pengfan Du, Yilin Niu +7

This study explores the scaling properties of Reinforcement Learning from Human Feedback (RLHF) in Large Language Models (LLMs). Although RLHF is considered an important step in po…

cs.CR2024

BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models

Xinyuan Wang, Victor Shea-Jay Huang, Renmiao Chen +4

While large language models (LLMs) exhibit remarkable capabilities across various tasks, they encounter potential security risks such as jailbreak attacks, which exploit vulnerabil…