4 citations · 10 across the 5 of their papers we have counts for
5 papers
SuperHF: Supervised Iterative Learning from Human Feedback
Gabriel Mukobi, Peter Chatain, Su Fong +4
While large language models demonstrate remarkable capabilities, they often present challenges in terms of safety, alignment with human values, and stability during training. Here,…
Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models
Mayee F. Chen, Nicholas Roberts, Kush Bhatia +4
The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good…
Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot Classification
Neel Guha, Mayee F. Chen, Kush Bhatia +3
Recent work has shown that language models' (LMs) prompt-based learning capabilities make them well suited for automating data labeling in domains where manual annotation is expens…
Reward Learning as Doubly Nonparametric Bandits: Optimal Design and Scaling Laws
Kush Bhatia, Wenshuo Guo, Jacob Steinhardt
Specifying reward functions for complex tasks like object manipulation or driving is challenging to do by hand. Reward learning seeks to address this by learning a reward model usi…
Congested Bandits: Optimal Routing via Short-term Resets
Pranjal Awasthi, Kush Bhatia, Sreenivas Gollapudi +1
For traffic routing platforms, the choice of which route to recommend to a user depends on the congestion on these routes -- indeed, an individual's utility depends on the number o…