2 papers
cs.CL2026
Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment
Yucong Huang, Xiucheng Li, Kaiqi Zhao +1
Standard RLHF relies on transitive scalar rewards, failing to capture the cyclic nature of human preferences. While some approaches like the General Preference Model (GPM) address…
cs.AI2026
CreativeGame:Toward Mechanic-Aware Creative Game Generation
Hongnan Ma, Han Wang, Shenglin Wang +6
Large language models can generate plausible game code, but turning this capability into \emph{iterative creative improvement} remains difficult. In practice, single-shot generatio…