From the 1 of 4 linked papers with an AI index.
4 papers
Transcoders for Investigating Deception in Language Models
Darius Lim, Nathan Leow, Xin Wei Chia
The paper uses per‑layer transcoders to build attribution graphs that reveal internal features linked to deceptive outputs in a Qwen3‑4B language model, showing how deception can b…
Multi-Trait Subspace Steering to Reveal the Dark Side of Human-AI Interaction
Xin Wei Chia, Swee Liang Wong, Jonathan Pan
Recent incidents have highlighted alarming cases where human-AI interactions led to negative psychological outcomes, including mental health crises and even user harm. As LLMs serv…
Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
Xin Wei Chia, Swee Liang Wong, Jonathan Pan
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt…
Prompt Inject Detection with Generative Explanation as an Investigative Tool
Jonathan Pan, Swee Liang Wong, Yidi Yuan +1
Large Language Models (LLMs) are vulnerable to adversarial prompt based injects. These injects could jailbreak or exploit vulnerabilities within these models with explicit prompt r…