2 citations · 2 across the 4 of their papers we have counts for
3 papers · 1 filter
Forecasting Side Effects of Activation Steering
Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang +1
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, ste…
NuHF Claw: A Risk Constrained Cognitive Agent Framework for Human Centered Procedure Support in Digital Nuclear Control Rooms
Xingyu Xiao, Jiejuan Tong, Jun Sun +4
The rapid digitization of nuclear power plant main control rooms has fundamentally reshaped operator interaction patterns, introducing complex soft-control behaviors and elevated c…
LLMScan: Causal Scan for LLM Misbehavior Detection
Mengdi Zhang, Kai Kiat Goh, Peixin Zhang +3
Despite the success of Large Language Models (LLMs) across various fields, their potential to generate untruthful, biased and harmful responses poses significant risks, particularl…