most citedMM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

7 citations · 15 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

Inference-Time Language Model Alignment via Integrated Value Guidance

Zhixuan Liu, Zhanhui Zhou, Yuanfu Wang +2

Large language models are typically fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. In this work, we introduce $\texti…

cs.CL2024

Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models

Zhanhui Zhou, Zhixuan Liu, Jie Liu +3

Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{we…

cs.CL2024★ 1 cited

Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Zhichen Dong, Zhanhui Zhou, Chao Yang +2

Large Language Models (LLMs) are now commonplace in conversation applications. However, their risks of misuse for generating harmful responses have raised serious societal concerns…

cs.CL2024

Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

Zhanhui Zhou, Jie Liu, Zhichen Dong +4

Large language models (LLMs) undergo safety alignment to ensure safe conversations with humans. However, this paper introduces a training-free attack method capable of reversing sa…

cs.CL2024★ 5 cited

SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning

Guoxin Chen, Kexin Tang, Chao Yang +3

Elucidating the reasoning process with structured explanations from question to answer is crucial, as it significantly enhances the interpretability, traceability, and trustworthin…