3 papers
cs.CR2026
Jailbreaking for the Average Jane: Choosing Optimal Jailbreaks via Bandit Algorithms for Automatically Enhanced Queries
Prarabdh Shukla, Ritik, Suhas Rao +2
With a profusion of jailbreaks for LLMs now widely known, a growing concern is that non-expert malicious actors ("the average Jane") could elicit actionable responses to malicious…
cs.CL2025
Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
Prarabdh Shukla, Wei Yin Chong, Yash Patel +3
To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement(, users commenting on live stream…
cs.LG2024
DiffRed: Dimensionality Reduction guided by stable rank
Prarabdh Shukla, Gagan Raj Gupta, Kunal Dutta
In this work, we propose a novel dimensionality reduction technique, DiffRed, which first projects the data matrix, A, along first principal components and the residual matri…