Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Toward Faithful and Complete Answer Construction from a Single Document
Zhaoyang Chen, Cody Fleming
Modern large language models (LLMs) are powerful generators driven by statistical next-token prediction. While effective at producing fluent text, this design biases models toward…
cs.LG2025
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
Zhaoyang Chen, Cody Fleming
Classifier free guidance has shown strong potential in diffusion-based reinforcement learning. However, existing methods rely on joint training of the guidance module and the diffu…