5 papers
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-intern…
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Test-time reinforcement learning (TTRL) reports substantial accuracy gains on mathematical reasoning benchmarks using majority vote as a pseudo-label signal. We argue these gains a…
When Relevance Meets Novelty: Dual-Stable Periodic Optimization for Serendipitous Recommendation
Hongxiang Lin, Hao Guo, Zeshun Li +6
Traditional recommendation systems tend to trap users in strong feedback loops by excessively pushing content aligned with their historical preferences, thereby limiting exploratio…
Dynamic Forgetting and Spatio-Temporal Periodic Interest Modeling for Local-Life Service Recommendation
Zhaoyu Hu, Jianyang Wang, Hao Guo +6
In the context of the booming digital economy, recommendation systems, as a key link connecting users and numerous services, face challenges in modeling user behavior sequences on…
Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation
Hao Guo, Erpeng Xue, Lei Huang +5
Deep Learning Recommendation Models (DLRMs) often rely on extensive manual feature engineering to improve accuracy and user experience, which increases system complexity and limits…