2 papers
cs.CL2026
CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning
Rahul Markasserithodi, Aditya Joshi, Yuekang Li +3
Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can bypass safety filters even on f…
cs.CL2025
Nek Minit: Harnessing Pragmatic Metacognitive Prompting for Explainable Sarcasm Detection of Australian and Indian English
Ishmanbir Singh, Dipankar Srirag, Aditya Joshi
Sarcasm is a challenge to sentiment analysis because of the incongruity between stated and implied sentiment. The challenge is exacerbated when the implication may be relevant to a…