Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Experience-Efficient Model-Free Deep Reinforcement Learning Using Pre-Training
Ruoxing Yang
We introduce PPOPT - Proximal Policy Optimization using Pretraining, a novel, model-free deep-reinforcement-learning algorithm that leverages pretraining to achieve high training e…
cs.LG2025
DP-Adam-AC: Privacy-preserving Fine-Tuning of Localizable Language Models Using Adam Optimization with Adaptive Clipping
Ruoxing Yang
Large language models (LLMs) such as ChatGPT have evolved into powerful and ubiquitous tools. Fine-tuning on small datasets allows LLMs to acquire specialized skills for specific t…