1 paper · 1 filter
Andrew Kiruluta, Andreas Lemos, Priscilla Burity
We propose a novel reinforcement learning framework for post training large language models that does not rely on human in the loop feedback. Instead, our approach uses cross atten…