8 citations · 15 across the 12 of their papers we have counts for
4 papers · 1 filter
When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs
Apoorva Upadhyaya, Sandipan Sikdar
Safety alignment of large language models (LLMs) degrades across languages, yet the internal mechanism driving this asymmetry remains poorly understood. Our work, therefore, presen…
ACCESS DENIED INC: The First Benchmark Environment for Sensitivity Awareness
Dren Fazlija, Arkadij Orlov, Sandipan Sikdar
Large language models (LLMs) are increasingly becoming valuable to corporate data management due to their ability to process text from various document formats and facilitate user…
SensePOLAR: Word sense aware interpretability for pre-trained contextual word embeddings
Jan Engler, Sandipan Sikdar, Marlene Lutz +1
Adding interpretability to word embeddings represents an area of active research in text representation. Recent work has explored thepotential of embedding words via so-called pola…
The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word Embeddings
Binny Mathew, Sandipan Sikdar, Florian Lemmerich +1
We introduce POLAR - a framework that adds interpretability to pre-trained word embeddings via the adoption of semantic differentials. Semantic differentials are a psychometric con…