3 citations · 3 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Two Is Better Than One: Rotations Scale LoRAs
Hongcan Guo, Guoshun Nan, Yuan Yang +9
Scaling Low-Rank Adaptation (LoRA)-based Mixture-of-Experts (MoE) facilitates large language models (LLMs) to efficiently adapt to diverse tasks. However, traditional gating mechan…
cs.LG2023★ 3 cited
Policy Regularization with Dataset Constraint for Offline Reinforcement Learning
Yuhang Ran, Yi-Chen Li, Fuxiang Zhang +2
We consider the problem of learning the best possible policy from a fixed dataset, known as offline Reinforcement Learning (RL). A common taxonomy of existing offline RL works is p…