Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
Ruiyi Ding, Yongxuan Lv, Xianhui Meng +4
Policy optimization for large language models often suffers from sparse reward signals in multi-step reasoning tasks. Critic-free methods like GRPO assign a single normalized outco…
cs.LG2025
Tracing the Heart's Pathways: ECG Representation Learning from a Cardiac Conduction Perspective
Tan Pan, Yixuan Sun, Chen Jiang +8
The multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects f…