3 papers
cs.AI2026
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Matthew Nguyen, Kyle Cox, Austin Meek +1
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is p…
cs.AI2026
Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers
Kyle Cox, Darius Kianersi, Adrià Garriga-Alonso
As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, sugges…
cs.CL2025
Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
Kyle Cox, Jiawei Xu, Yikun Han +6
An interesting behavior in large language models (LLMs) is prompt sensitivity. When provided with different but semantically equivalent versions of the same prompt, models may prod…