80 citations · 132 across the 22 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
Donglai Xu, Hongzheng Yang, Yuzhi Zhao +10
Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models (MLLMs) is highly dependent on high-quality labeled data, which is often scarce and prone…
cs.LG2024
FedRepOpt: Gradient Re-parametrized Optimizers in Federated Learning
Kin Wai Lau, Yasar Abbas Ur Rehman, Pedro Porto Buarque de Gusmão +3
Federated Learning (FL) has emerged as a privacy-preserving method for training machine learning models in a distributed manner on edge devices. However, on-device models face inhe…