Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
Ziru Chen, Dongdong Chen, Ruinan Jin +3
Recently, there have been significant research interests in training large language models (LLMs) with reinforcement learning (RL) on real-world tasks, such as multi-turn code gene…
cs.LG2025
Rethinking the Vulnerability of Concept Erasure and a New Method
Alex D. Richardson, Kaicheng Zhang, Lucas Beerens +1
The proliferation of text-to-image diffusion models has raised significant privacy and security concerns, particularly regarding the generation of copyrighted or harmful images. In…