inference manipulation 1instruction following 1large language models 1model safety 1prompt hierarchy 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
Steering Instruction Hierarchies at Inference Time
Siqi Zeng, Sewoong Lee, Han Zhao +1
The paper proposes V‑Steer, a training‑free method that modifies cached value vectors during inference to ensure higher‑priority prompts (like system prompts) override lower‑priori…
cs.CL2026
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…