1 paper
Karthika Arumugam, Kiran Kumar Manku, Amit Dhanda
Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these agents condition on natural langu…