2 papers
cs.LG2024
Scores as Actions: a framework of fine-tuning diffusion models by continuous-time reinforcement learning
Hanyang Zhao, Haoxian Chen, Ji Zhang +2
Reinforcement Learning from human feedback (RLHF) has been shown a promising direction for aligning generative models with human intent and has also been explored in recent works f…
cs.CL2022
Incorporating Causal Analysis into Diversified and Logical Response Generation
Jiayi Liu, Wei Wei, Zhixuan Chu +4
Although the Conditional Variational AutoEncoder (CVAE) model can generate more diversified responses than the traditional Seq2Seq model, the responses often have low relevance wit…