Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Infusion: Shaping Model Behavior by Editing Training Data via Influence Functions
J Rosser, Robert Kirk, Edward Grefenstette +2
Influence functions are commonly used to attribute model behavior to training documents. We explore the reverse: crafting training data that induces model behavior. Our framework,…
cs.LG2025
Mapping Faithful Reasoning in Language Models
Jiazheng Li, Andreas Damianou, J Rosser +2
Chain-of-thought (CoT) traces promise transparency for reasoning language models, but prior work shows they are not always faithful reflections of internal computation. This raises…