3 papers
cs.AI2026
Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
Adhyyan Narang, Artin Tajdini, Claire Zhang +1
Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferen…
cs.LG2026
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
Adhyyan Narang, Sarah Dean, Lillian J Ratliff +1
In many economically relevant contexts where machine learning is deployed, multiple platforms obtain data from the same pool of users, each of whom selects the platform that best s…
cs.LG2025
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
Marcus Williams, Micah Carroll, Adhyyan Narang +3
As LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators.…