activation steering 1evaluation awareness 1large language models 1model safety 1prompt optimization 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1
The paper investigates how to suppress specific internal activations in large language models by optimizing only the input prompt, aiming to hide evaluation-awareness signals witho…
cs.CL2026
Joint Optimization for Greedy Longest-match Tokenization
Adhiraj Singh, Deepanshu Mody, Ghina Al Shdaifat +4
Recent work has shown that subword vocabularies can be trained to optimize compression for a specific inference rule rather than relying on greedy heuristics such as Byte Pair Enco…