1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 1 cited
Critique-out-Loud Reward Models
Zachary Ankner, Mansheej Paul, Brandon Cui +2
Traditionally, reward models used for reinforcement learning from human feedback (RLHF) are trained to directly predict preference scores without leveraging the generation capabili…
cs.LG2023★ 1 cited
Striped Attention: Faster Ring Attention for Causal Transformers
William Brandon, Aniruddha Nrusimha, Kevin Qian +4
To help address the growing demand for ever-longer sequence lengths in transformer models, Liu et al. recently proposed Ring Attention, an exact attention algorithm capable of over…