3 citations · 3 across the 2 of their papers we have counts for
3 papers
When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs
Apoorva Upadhyaya, Sandipan Sikdar
Safety alignment of large language models (LLMs) degrades across languages, yet the internal mechanism driving this asymmetry remains poorly understood. Our work, therefore, presen…
Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability
Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella
Zero-Shot Stance Detection (ZSSD) identifies the attitude of the post toward unseen targets. Existing research using contrastive, meta-learning, or data augmentation suffers from g…
A Multi-task Model for Sentiment Aided Stance Detection of Climate Change Tweets
Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl
Climate change has become one of the biggest challenges of our time. Social media platforms such as Twitter play an important role in raising public awareness and spreading knowled…