5 papers
Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
Adhyyan Narang, Artin Tajdini, Claire Zhang +1
Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferen…
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
Artin Tajdini, Jonathan Scarlett, Kevin Jamieson
We study stochastic linear bandits with heavy-tailed rewards, where the rewards have a finite -absolute central moment bounded by for some . We impr…
Learning to Incentivize in Repeated Principal-Agent Problems with Adversarial Agent Arrivals
Junyan Liu, Arnab Maiti, Artin Tajdini +2
We initiate the study of a repeated principal-agent problem over a finite horizon , where a principal sequentially interacts with types of agents arriving in an advers…
Nearly Minimax Optimal Submodular Maximization with Bandit Feedback
Artin Tajdini, Lalit Jain, Kevin Jamieson
We consider maximizing an unknown monotonic, submodular set function with cardinality constraint under stochastic bandit feedback. At each time $t=1,…
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
Haolin Liu, Artin Tajdini, Andrew Wagenmaker +1
In linear bandits, how can a learner effectively learn when facing corrupted rewards? While significant work has explored this question, a holistic understanding across different a…