2 papers
cs.LG2026
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
Shangpin Peng, Weinong Wang, Zhuotao Tian +7
Direct Preference Optimization (DPO) has emerged as a cornerstone of reinforcement learning from human feedback (RLHF) due to its simplicity and efficiency. However, existing DPO-b…
nucl-th2026
Impact of the in-medium cross section on cluster spectra in collisions at and
C. K. Tam, Z. Chajecki, R. S. Wang +35
Although significant efforts have been made to investigate the density dependence of the nuclear symmetry energy, the influence of the in-medium cross section on particle productio…