#reward optimization
try —
2 papers match
cs.LG2026
HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
Tiangang Li, Xiangbo Tian
The paper introduces HARGO, a reinforcement‑learning post‑training method that weights responses by confidence and reward contrast to better align large language models with divers…
#large language models#reinforcement learning#high-performance computing#reward optimization
cs.CV2026
MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
Yinhan Zhang, Dingwei Tan, Dinwei Tan +4
The paper introduces MagicPrompt, a lightweight method that uses attention-embedded soft prompts and dual-space reward feedback to fine‑tune large video diffusion models with less…
#video generation#diffusion models#prompt tuning#parameter-efficient fine-tuning