1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2026
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning
Xinyu Ma, Mingzhou Xu, Xuebo Liu +4
Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore…
cs.IR2025★ 1 cited
OpenOneRec Technical Report
Guorui Zhou, Honghui Bao, Jiaming Huang +44
While the OneRec series has successfully unified the fragmented recommendation pipeline into an end-to-end generative framework, a significant gap remains between recommendation sy…