activity
20182024
most citedIntermediate-Task Transfer Learning with Pretrained Models for Natural Language Understanding: When and Why Does It Work?

53 citations · 195 across the 20 of their papers we have counts for

collaborators

26 papers

cs.CL2024

Transformers Struggle to Learn to Search

Abulhair Saparov, Srushti Pawar, Shreyas Pimpalgaonkar +6

Search is an ability foundational in many important tasks, and recent studies have shown that large language models (LLMs) struggle to perform search robustly. It is unknown whethe…

cs.CL2024

Self-Generated Critiques Boost Reward Modeling for Language Models

Yue Yu, Zhengxing Chen, Aston Zhang +10

Reward modeling is crucial for aligning large language models (LLMs) with human preferences, especially in reinforcement learning from human feedback (RLHF). However, current rewar…

cs.CL2024

Self-Consistency Preference Optimization

Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang +6

Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex…

cs.CL2024★ 3 cited

Self-Taught Evaluators

Tianlu Wang, Ilia Kulikov, Olga Golovneva +7

Model-based evaluation is at the heart of successful model development -- as a reward model for training, and as a replacement for human evaluation. To train such evaluators, the s…

cs.LG2024★ 37 cited

An Introduction to Vision-Language Modeling

Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay +38

Following the recent popularity of Large Language Models (LLMs), several attempts have been made to extend them to the visual domain. From having a visual assistant that could guid…

cs.CL2024★ 4 cited

Iterative Reasoning Preference Optimization

Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3

Iterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks (Y…