1 citations · 1 across the 10 of their papers we have counts for
11 papers
Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
Zhongyi Li, Wan Tian, Jingyu Chen +8
Multi-agent collaboration has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models, yet it suffers from interaction-level ambiguity that…
Your Group-Relative Advantage Is Biased
Fengkai Yang, Zherui Chen, Xiaohan Wang +10
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…
LLMBoost: Make Large Language Models Stronger with Boosting
Zehao Chen, Tianxiang Ai, Yifei Li +11
Ensemble learning of LLMs has emerged as a promising alternative to enhance performance, but existing approaches typically treat models as black boxes, combining the inputs or fina…
FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
Zheng Dou, Deqing Wang, Fuzhen Zhuang +2
Scientific document representation learning provides powerful embeddings for various tasks, while current methods face challenges across three approaches. 1) Contrastive training w…
Data-Free Continual Learning of Server Models in Model-Heterogeneous Cloud-Device Collaboration
Xiao Zhang, Zengzhe Chen, Yuan Yuan +5
The rise of cloud-device collaborative computing has enabled intelligent services to be delivered across distributed edge devices while leveraging centralized cloud resources. In t…
ORCA: Mitigating Over-Reliance for Multi-Task Dwell Time Prediction with Causal Decoupling
Huishi Luo, Fuzhen Zhuang, Yongchun Zhu +6
Dwell time (DT) is a critical post-click metric for evaluating user preference in recommender systems, complementing the traditional click-through rate (CTR). Although multi-task l…