3 papers
cs.AI2026
FaithSteer-BENCH: A Deployment-Aligned Stress-Testing Benchmark for Inference-Time Steering
Zikang Ding, Qiying Hu, Yi Zhang +4
Inference-time steering is widely regarded as a lightweight and parameter-free mechanism for controlling large language model (LLM) behavior, and prior work has often suggested tha…
cs.CL2025
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
Hongji Li, Junchi yao, Manjiang Yu +4
Machine unlearning aims to erase requested data from trained models without full retraining. For Reasoning Multimodal Large Language Models (RMLLMs), this is uniquely challenging:…
cs.AI2025
PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
Manjiang Yu, Hongji Li, Priyanka Singh +3
Reliable behavior control is central to deploying large language models (LLMs) on the web. Activation steering offers a tuning-free route to align attributes (e.g., truthfulness) t…