4 papers
GIPO: Gaussian Importance Sampling Policy Optimization
Chengxuan Lu, Zhenquan Zhang, Shukuan Wang +3
Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor da…
AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models
Chengxuan Lu, Shukuan Wang, Yanjie Li +10
Reinforcement learning (RL) for large-scale Vision-Language-Action (VLA) models is severely bottlenecked by synchronization barriers and the high cost of environment data acquisiti…
DUET: Dual Model Co-Training for Entire Space CTR Prediction
Yutian Xiao, Meng Yuan, Fuzhen Zhuang +9
The pre-ranking stage plays a pivotal role in large-scale recommender systems but faces an intrinsic trade-off between model expressiveness and computational efficiency. Owing to t…
MARS: Modality-Aligned Retrieval for Sequence Augmented CTR Prediction
Yutian Xiao, Shukuan Wang, Binhao Wang +6
Click-through rate (CTR) prediction serves as a cornerstone of recommender systems. Despite the strong performance of current CTR models based on user behavior modeling, they are s…