1 paper
Chirag Nagpal, Subhashini Venugopalan, Jimmy Tobin +3
We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to a…