2 papers
cs.CR2026
Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks
Md Rysul Kabir, Zoran Tiganj
Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different fai…
cs.AI2024
Deep reinforcement learning with time-scale invariant memory
Md Rysul Kabir, James Mochizuki-Freeman, Zoran Tiganj
The ability to estimate temporal relationships is critical for both animals and artificial agents. Cognitive science and neuroscience provide remarkable insights into behavioral an…