Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
Siyi Gu, Jialin Chen, Sophia Zhou +2
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillation often relies on chain-of-t…
cs.AI2024
Thought Propagation: An Analogical Approach to Complex Reasoning with Large Language Models
Junchi Yu, Ran He, Rex Ying
Large Language Models (LLMs) have achieved remarkable success in reasoning tasks with the development of prompting methods. However, existing prompting approaches cannot reuse insi…