2 papers
cs.CR2026
Jailbreaking for the Average Jane: Choosing Optimal Jailbreaks via Bandit Algorithms for Automatically Enhanced Queries
Prarabdh Shukla, Ritik, Suhas Rao +2
With a profusion of jailbreaks for LLMs now widely known, a growing concern is that non-expert malicious actors ("the average Jane") could elicit actionable responses to malicious…
cs.LG2026
Creator Incentives in Recommender Systems: A Cooperative Game-Theoretic Approach for Stable and Fair Collaboration in Multi-Agent Bandits
Ramakrishnan Krishnamurthy, Arpit Agarwal, Lakshminarayanan Subramanian +1
User interactions in online recommendation platforms create interdependencies among content creators: feedback on one creator's content influences the system's learning and, in tur…