activity
20232026
most citedEthical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

3 citations · 4 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2026

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

Abhinav Rao, Liancheng Gong, Bin Hu +1

Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, a…

cs.CL2024

[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs

Abhinav Rao, Monojit Choudhury, Somak Aditya

We introduce two paradoxes concerning jailbreak of foundation models: First, it is impossible to construct a perfect jailbreak classifier, and second, a weaker model cannot consist…

cs.CL2024

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

Abhinav Rao, Akhila Yerukola, Vishwa Shah +2

To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We i…

cs.CL20233 cited

Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

Abhinav Rao, Aditi Khandelwal, Kumar Tanmay +2

In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so t…

cs.CR20231 cited

MALITE: Lightweight Malware Detection and Classification for Constrained Devices

Sidharth Anand, Barsha Mitra, Soumyadeep Dey +3

Today, malware is one of the primary cyberthreats to organizations. Malware has pervaded almost every type of computing device including the ones having limited memory, battery and…