activity
20212024
most citedTraining Value-Aligned Reinforcement Learning Agents Using a Normative Prior

6 citations · 26 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2024

A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds

Christopher Z. Cui, Xiangyu Peng, Mark O. Riedl

Open-ended worlds are those in which there are no pre-specified goals or environmental reward signal. As a consequence, an agent must know how to perform a multitude of tasks. Howe…

cs.CL20234 cited

Dialogue Shaping: Empowering Agents through NPC Interaction

Wei Zhou, Xiangyu Peng, Mark Riedl

One major challenge in reinforcement learning (RL) is the large amount of steps for the RL agent needs to converge in the training process and learn the optimal policy, especially…

cs.CL2023

Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models

Ashutosh Baheti, Ximing Lu, Faeze Brahman +3

Reinforcement Learning with Human Feedback (RLHF) is the most prominent method for Language Model (LM) alignment. However, RLHF is an unstable and data-hungry process that continua…

cs.CL2023

Few-Shot Dialogue Summarization via Skeleton-Assisted Prompt Transfer in Prompt Tuning

Kaige Xie, Tong Yu, Haoliang Wang +6

In real-world scenarios, labeled samples for dialogue summarization are usually limited (i.e., few-shot) due to high annotation costs for high-quality dialogue summaries. To effici…

cs.CL2021

Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts

Ashutosh Baheti, Maarten Sap, Alan Ritter +1

Dialogue models trained on human conversations inadvertently learn to generate toxic responses. In addition to producing explicitly offensive utterances, these models can also impl…

cs.CL2021

Fabula Entropy Indexing: Objective Measures of Story Coherence

Louis Castricato, Spencer Frazier, Jonathan Balloch +1

Automated story generation remains a difficult area of research because it lacks strong objective measures. Generated stories may be linguistically sound, but in many cases suffer…