5 citations · 5 across the 5 of their papers we have counts for
4 papers · 1 filter
Triaging Threats to Specialized Guardrails
Wenjie Jacky Mo, Xiaofei Wen, Rui Cai +6
Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains challenging because safety ris…
When AI Meets Wall Street: A Survey on Trustworthy AI in Fintech
Qingwen Zeng, Zhenghao Zhao, Yitian Yang +6
Artificial intelligence is now embedded as a primary decision engine in continuously operated financial AI pipelines spanning training and updating, deployment and inference, and o…
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
Yingjie Zhang, Tong Liu, Zhe Zhao +2
LLMs remain vulnerable to jailbreak attacks that exploit adversarial prompts to circumvent safety measures. Current safety fine-tuning approaches face two critical limitations. Fir…
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Tong Liu, Yingjie Zhang, Zhe Zhao +3
In recent years, large language models (LLMs) have demonstrated notable success across various tasks, but the trustworthiness of LLMs is still an open problem. One specific threat…