2 citations · 3 across the 5 of their papers we have counts for
1 paper · 1 filter
Harsh Kumar, Rahul Maity, Tanmay Joshi +4
Aligned large language models (LLMs) remain vulnerable to adversarial manipulation, and their reliance on web-scale pretraining creates a subtle but consequential attack surface. W…