Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
Wei Deng
We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.g., GRPO) does not reach. We fin…
cs.CL2026
Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space
ZiYi Dong, Yuliang Huang, Weijian Deng +3
This work reformulates language generation as a stochastic optimal control problem, providing a unified theoretical perspective to analyze autoregressive and diffusion models and e…