5 papers
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-intern…
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Test-time reinforcement learning (TTRL) reports substantial accuracy gains on mathematical reasoning benchmarks using majority vote as a pseudo-label signal. We argue these gains a…
When Relevance Meets Novelty: Dual-Stable Periodic Optimization for Serendipitous Recommendation
Hongxiang Lin, Hao Guo, Zeshun Li +6
Traditional recommendation systems tend to trap users in strong feedback loops by excessively pushing content aligned with their historical preferences, thereby limiting exploratio…
Dynamic Forgetting and Spatio-Temporal Periodic Interest Modeling for Local-Life Service Recommendation
Zhaoyu Hu, Jianyang Wang, Hao Guo +6
In the context of the booming digital economy, recommendation systems, as a key link connecting users and numerous services, face challenges in modeling user behavior sequences on…
MTmixAtt: Integrating Mixture-of-Experts with Multi-Mix Attention for Large-Scale Recommendation
Xianyang Qi, Yuan Tian, Zhaoyu Hu +4
Industrial recommender systems critically depend on high-quality ranking models. However, traditional pipelines still rely on manual feature engineering and scenario-specific archi…