4 papers
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
Yiqi Zhang, Fangzheng Jiao, Tian Tang +13
Reinforcement learning with verifiable rewards (RLVR) has recently unlocked strong reasoning capabilities in large language models (LLMs), triggering rapid exploration of new algor…
Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
Lin Guan, Jia-Qi Yang, Zhishan Zhao +12
Short-video recommenders such as Douyin must exploit extremely long user behavior histories without breaking latency or cost budgets. We present an end-to-end industrial recommende…
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
Pan Wang, Yihao Hu, Xiujin Liu +3
Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most existing frameworks store me…
RankMixer: Scaling Up Ranking Models in Industrial Recommenders
Jie Zhu, Zhifang Fan, Xiaoxie Zhu +18
Recent progress on large language models (LLMs) has spurred interest in scaling up recommendation systems, yet two practical obstacles remain. First, training and serving cost on i…