2 papers
cs.GT2025
Learning in Stackelberg Games with Non-myopic Agents
Nika Haghtalab, Thodoris Lykouris, Sloan Nietert +1
We study Stackelberg games where a principal repeatedly interacts with a non-myopic long-lived agent, without knowing the agent's payoff function. Although learning in Stackelberg…
cs.CR2024
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
Danny Halawi, Alexander Wei, Eric Wallace +3
Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety…