3 papers
cs.RO2026
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
Jihoon Hong, Julian Skifstad, Qiyue Dai +2
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study…
cs.LG2026
Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control
Jihoon Hong, Alice Chan, Qiyue Dai +2
Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs without sacrificing visual quality.…
cs.LG2026
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
Julian Skifstad, Xinyue Annie Yang, Glen Chou
Inference-time LLM alignment methods, particularly activation steering, offer an alternative to fine-tuning by directly modifying activations during generation. Existing methods, h…