1 citations · 1 across the 5 of their papers we have counts for
3 papers · 1 filter
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
Jiajia Tang, Sizhe Yuen, Francisco Gomez Medina +2
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interf…
Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?
Hao Liang, Jiayu Cheng, Sean R. Sinclair +1
Exogenous MDPs (Exo-MDPs) capture sequential decision-making where uncertainty comes solely from exogenous inputs that evolve independently of the learner's actions. This structure…
Reduced Policy Optimization for Continuous Control with Hard Constraints
Shutong Ding, Jingya Wang, Yali Du +1
Recent advances in constrained reinforcement learning (RL) have endowed reinforcement learning with certain safety guarantees. However, deploying existing constrained RL algorithms…