1 paper
Ihor Vitenko, Noha Ibrahim, Sihem Amer-Yahia
Reinforcement learning (RL) policies are typically trained for fixed objectives, making reuse difficult when task requirements change. We study inference-time policy reuse: given a…