Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
Darpan Aswal, Céline Hudelot
Large Language Models have found success in a variety of applications. However, their safety remains a concern due to the existence of various jailbreaking methods. Despite signifi…
cs.CL2025
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
Darpan Aswal, Siddharth D Jaiswal
Safety-aligned LLMs remain vulnerable to digital phenomena like textese that introduce non-canonical perturbations to words but preserve the phonetics. We introduce CMP-RT (code-mi…
cs.CL2025
Efficient Environmental Claim Detection with Hyperbolic Graph Neural Networks
Darpan Aswal, Manjira Sinha
Transformer based models, especially large language models (LLMs) dominate the field of NLP with their mass adoption in tasks such as text generation, summarization and fake news d…