2 papers
cs.CL2026
Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time
Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot +5
Decode-time alignment methods steer a frozen language model by scoring candidate continuations with an external reward and selecting the maximiser. We argue that this shared design…
cs.CL2026
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
Partha Pratim Saha, Samarth Raina, Mayur Parvatikar +4
Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally. Aligned systems still fa…