8 citations · 11 across the 9 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement
Shuxing Yang, Kaihao Zhu, Junjie Yang +13
Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous r…
cs.LG2026
From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement Learning
Jun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen +1
Preference-based reinforcement learning (PbRL) avoids explicit reward engineering by learning from pairwise human preference feedback. Existing offline PbRL methods typically follo…