47 citations · 47 across the 2 of their papers we have counts for
1 paper · 1 filter
Hakim Sidahmed, Samrat Phatale, Alex Hutcheson +16
While Reinforcement Learning from Human Feedback (RLHF) effectively aligns pretrained Large Language and Vision-Language Models (LLMs, and VLMs) with human preferences, its computa…