activity
20242026
collaborators

10 papers

cs.AI2026

MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer

Ajsal Shereef Palattuparambil, Thommen George Karimpanal, Santu Rana

Humans excel at analogical reasoning - applying knowledge from one task to a related one with minimal relearning. In contrast, reinforcement learning (RL) agents typically require…

cs.AI2026

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4

Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show th…

cs.LG2026

Leveraging Human Feedback for Semantically-Relevant Skill Discovery

Maxence Hussonnois, Thommen George Karimpanal, Santu Rana

Unsupervised skill discovery in reinforcement learning aims to intrinsically motivate agents to discover diverse and useful behaviours. However, unconstrained approaches can produc…

cs.AI2026

ASPECT:Analogical Semantic Policy Execution via Language Conditioned Transfer

Ajsal Shereef Palattuparambil, Thommen George Karimpanal, Santu Rana

Reinforcement Learning (RL) agents often struggle to generalize knowledge to new tasks, even those structurally similar to ones they have mastered. Although recent approaches have…

cs.CL2026

The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs

Omar Mahmoud, Ali Khalil, Buddhika Laknath Semage +2

Hallucination in large language models (LLMs) has been widely studied in recent years, with progress in both detection and mitigation aimed at improving truthfulness. Yet, a critic…

cs.AI2025

Dynamic Policy Fusion for User Alignment Without Re-Interaction

Ajsal Shereef Palattuparambil, Thommen George Karimpanal, Santu Rana

Deep reinforcement learning (RL) policies, although optimal in terms of task rewards, may not align with the personal preferences of human users. To ensure this alignment, a naive…