6 papers
Assessing Domain-Level Susceptibility to Emergent Misalignment from Narrow Finetuning
Abhishek Mishra, Mugilan Arulvanan, Reshma Ashok +3
Emergent misalignment poses risks to AI safety as language models are increasingly used for autonomous tasks. In this paper, we present a population of large language models (LLMs)…
Exploring Re-inforcement Learning via Human Feedback under User Heterogeneity
Sarvesh Shashidhar, Abhishek Mishra, Madhav Kotecha
Re-inforcement learning from human feedback (RLHF) has been effective in the task of AI alignment. However, one of the key assumptions of RLHF is that the annotators (referred to a…
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
Arnav Ramamoorthy, Shrey Dhorajiya, Ojas Pungalia +6
Envy shapes competitiveness and cooperation in human groups, yet its role in large language model interactions remains largely unexplored. As LLMs increasingly operate in multi-age…
The Model's Language Matters: A Comparative Privacy Analysis of LLMs
Abhishek K. Mishra, Antoine Boutet, Lucas Magnana
Large Language Models (LLMs) are increasingly deployed across multilingual applications that handle sensitive data, yet their scale and linguistic variability introduce major priva…
Language translation, and change of accent for speech-to-speech task using diffusion model
Abhishek Mishra, Ritesh Sur Chowdhury, Vartul Bahuguna +2
Speech-to-speech translation (S2ST) aims to convert spoken input in one language to spoken output in another, typically focusing on either language translation or accent adaptation…
Guardians of Generation: Dynamic Inference-Time Copyright Shielding with Adaptive Guidance for AI Image Generation
Soham Roy, Abhishek Mishra, Shirish Karande +1
Modern text-to-image generative models can inadvertently reproduce copyrighted content memorized in their training data, raising serious concerns about potential copyright infringe…