most citedDistillSeq: A Framework for Safety Alignment Testing in Large Language Models using Knowledge Distillation

11 citations · 11 across the 3 of their papers we have counts for

collaborators

8 papers

cs.CR2025

Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models

Junzhe Yu, Yi Liu, Huijia Sun +2

Large Language Models (LLMs) have significantly advanced text understanding and generation, becoming integral to applications across education, software development, healthcare, en…

cs.CR2024

Efficient Detection of Toxic Prompts in Large Language Models

Yi Liu, Junzhe Yu, Huijia Sun +4

Large language models (LLMs) like ChatGPT and Gemini have significantly advanced natural language processing, enabling various applications such as chatbots and automated content g…

cs.CR2024

Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models

Zihao Xu, Yi Liu, Gelei Deng +4

Security concerns for large language models (LLMs) have recently escalated, focusing on thwarting jailbreaking attempts in discrete prompts. However, the exploration of jailbreak v…

cs.SE202411 cited

DistillSeq: A Framework for Safety Alignment Testing in Large Language Models using Knowledge Distillation

Mingke Yang, Yuqi Chen, Yi Liu +1

Large Language Models (LLMs) have showcased their remarkable capabilities in diverse domains, encompassing natural language understanding, translation, and even code generation. Th…

cs.CR2024

Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment

Jie Li, Yi Liu, Chongyang Liu +4

Large Language Models (LLMs) like OpenAI's GPT series, Anthropic's Claude, and Meta's LLaMa have shown remarkable capabilities in text generation. However, their susceptibility to…

cs.CR2024

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

Yuxi Li, Yi Liu, Yuekang Li +4

Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Current safety alignment techniques, how…