4 papers
Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models
Darpan Aswal, Thomas Palmeira Ferraz, Yongxin Zhou +1
Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decod…
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
Darpan Aswal, Siddharth D Jaiswal
Safety-aligned LLMs remain vulnerable to digital phenomena like textese that introduce non-canonical perturbations to words but preserve the phonetics. We introduce CMP-RT (code-mi…
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
Darpan Aswal, Céline Hudelot
Large Language Models have found success in a variety of applications. However, their safety remains a concern due to the existence of various jailbreaking methods. Despite signifi…
Efficient Environmental Claim Detection with Hyperbolic Graph Neural Networks
Darpan Aswal, Manjira Sinha
Transformer based models, especially large language models (LLMs) dominate the field of NLP with their mass adoption in tasks such as text generation, summarization and fake news d…