1 citations · 1 across the 4 of their papers we have counts for
4 papers
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
Weidi Luo, Shenghong Dai, Xiaogeng Liu +4
The rapid advancements in Large Language Models (LLMs) have enabled their deployment as autonomous agents for handling complex tasks in dynamic environments. These LLMs demonstrate…
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
Qin Liu, Fei Wang, Chaowei Xiao +1
The emergence of vision language models (VLMs) comes with increased safety concerns, as the incorporation of multiple modalities heightens vulnerability to attacks. Although VLMs c…
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
Qin Liu, Wenjie Mo, Terry Tong +4
The advancement of Large Language Models (LLMs) has significantly impacted various domains, including Web search, healthcare, and software development. However, as these models sca…
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
Xuefeng Du, Chaowei Xiao, Yixuan Li
The surge in applications of large language models (LLMs) has prompted concerns about the generation of misleading or fabricated information, known as hallucinations. Therefore, de…