2 papers
cs.AI2026
Multi-Trait Subspace Steering to Reveal the Dark Side of Human-AI Interaction
Xin Wei Chia, Swee Liang Wong, Jonathan Pan
Recent incidents have highlighted alarming cases where human-AI interactions led to negative psychological outcomes, including mental health crises and even user harm. As LLMs serv…
cs.LG2025
Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
Xin Wei Chia, Swee Liang Wong, Jonathan Pan
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt…