2 papers
cs.CL2025
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
Yidong Wang, Xin Wang, Cunxiang Wang +9
Self-Rewarding Language Models propose an architecture in which the Large Language Models(LLMs) both generates responses and evaluates its own outputs via LLM-as-a-Judge prompting,…
stat.ME2025
Doubly Robust Fusion of Many Treatments for Policy Learning
Ke Zhu, Jianing Chu, Ilya Lipkovich +2
Individualized treatment rules/recommendations (ITRs) aim to improve patient outcomes by tailoring treatments to the characteristics of each individual. However, when there are man…