3 citations · 4 across the 4 of their papers we have counts for
5 papers
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?
Abhinav Rao, Liancheng Gong, Bin Hu +1
Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, a…
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
Abhinav Rao, Monojit Choudhury, Somak Aditya
We introduce two paradoxes concerning jailbreak of foundation models: First, it is impossible to construct a perfect jailbreak classifier, and second, a weaker model cannot consist…
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
Abhinav Rao, Akhila Yerukola, Vishwa Shah +2
To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We i…
Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs
Abhinav Rao, Aditi Khandelwal, Kumar Tanmay +2
In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so t…
MALITE: Lightweight Malware Detection and Classification for Constrained Devices
Sidharth Anand, Barsha Mitra, Soumyadeep Dey +3
Today, malware is one of the primary cyberthreats to organizations. Malware has pervaded almost every type of computing device including the ones having limited memory, battery and…