advantage shaping 1contrastive learning 1on-policy distillation 1policy optimization 1token-level correctness 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Weiwen Xu, Jia Liu, Hou Pong Chan +4
The paper proposes Contrastive Policy Optimization, which leverages token‑level contrastive disagreement between reference‑guided and standard generation distributions to provide a…
cs.CL2024
Natural Language Fine-Tuning
Jia Liu, Yue Wang, Zhiqi Lin +3
Large language model fine-tuning techniques typically depend on extensive labeled data, external guidance, and feedback, such as human alignment, scalar rewards, and demonstration.…
cs.HC2024
FaGeL: Fabric LLMs Agent empowered Embodied Intelligence Evolution with Autonomous Human-Machine Collaboration
Jia Liu, Min Chen
Recent advancements in Large Language Models (LLMs) have enhanced the reasoning capabilities of embodied agents, driving progress toward AGI-powered robotics. While LLMs have been…