2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024★ 2 cited
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
Haochen Zhao, Xiangru Tang, Ziran Yang +8
The advancement and extensive application of large language models (LLMs) have been remarkable, including their use in scientific research assistance. However, these models often g…
cs.CL2024
Panacea: Pareto Alignment via Preference Adaptation for LLMs
Yifan Zhong, Chengdong Ma, Xiaoyuan Zhang +5
Current methods for large language model alignment typically use scalar human preference labels. However, this convention tends to oversimplify the multi-dimensional and heterogene…
cs.CL2023
Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
Chengdong Ma, Ziran Yang, Hai Ci +4
The primary challenge in deploying Large Language Model (LLM) is ensuring its harmlessness. Red team can identify vulnerabilities by attacking LLM to attain safety. However, curren…