papers

Publications (16)

cs.LG2021

Not All Memories are Created Equal: Learning to Forget by Expiring

Sainbayar Sukhbaatar, Da Ju, Spencer Poff +4

Attention mechanisms have shown promising results in sequence modeling tasks that require long-term memory. Recent work investigated mechanisms to reduce the computational cost of…

cs.LG2018

High-Level Strategy Selection under Partial Observability in StarCraft: Brood War

Jonas Gehring, Da Ju, Vegard Mella +3

We consider the problem of high-level strategy selection in the adversarial setting of real-time strategy games from a reinforcement learning perspective, where taking an action co…

cs.LG2021

Staircase Attention for Recurrent Processing of Sequences

Da Ju, Stephen Roller, Sainbayar Sukhbaatar +1

Attention mechanisms have become a standard tool for sequence modeling tasks, in particular by stacking self-attention layers over the entire input sequence as in the Transformer a…

cs.CL2020

All-in-One Image-Grounded Conversational Agents

Da Ju, Kurt Shuster, Y-Lan Boureau +1

As single-task accuracy on individual language and image tasks has improved substantially in the last few years, the long-term goal of a generally skilled agent that can both see a…

cs.CL2021

The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

Naman Goyal, Cynthia Gao, Vishrav Chaudhary +7

One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks eithe…

cs.CL2020

Recipes for building an open-domain chatbot

Stephen Roller, Emily Dinan, Naman Goyal +9

Building open-domain chatbots is a challenging area for machine learning research. While prior work has shown that scaling neural models in the number of parameters and the size of…

cs.CL2022

BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage

Kurt Shuster, Jing Xu, Mojtaba Komeili +15

We present BlenderBot 3, a 175B parameter dialogue model capable of open-domain conversation with access to the internet and a long-term memory, and having been trained on a large…

cs.CL2020

The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational Agents

Kurt Shuster, Da Ju, Stephen Roller +3

We introduce dodecaDialogue: a set of 12 tasks that measures if a conversational agent can communicate engagingly with personality and empathy, ask questions, answer questions by u…

cs.CL2021

Recipes for Safety in Open-domain Chatbots

Jing Xu, Da Ju, Margaret Li +3

Models trained on large unlabeled corpora of human interactions will learn patterns and mimic behaviors therein, which include offensive or otherwise toxic behavior and unwanted bi…

cs.CL2023

Improving Open Language Models by Learning from Organic Interactions

Jing Xu, Da Ju, Joshua Lane +10

We present BlenderBot 3x, an update on the conversational model BlenderBot 3, which is now trained using organic conversation and feedback data from participating users of the syst…

cs.CL2020

Multi-Modal Open-Domain Dialogue

Kurt Shuster, Eric Michael Smith, Da Ju +1

Recent work in open-domain conversational agents has demonstrated that significant improvements in model engagingness and humanness metrics can be achieved via massive scaling in b…

cs.CL2024

Are Female Carpenters like Blue Bananas? A Corpus Investigation of Occupation Gender Typicality

Da Ju, Karen Ulrich, Adina Williams

People tend to use language to mention surprising properties of events: for example, when a banana is blue, we are more likely to mention color than when it is yellow. This fact is…

cs.CY2024

Sense and Sensitivity: Evaluating the simulation of social dynamics via Large Language Models

Da Ju, Adina Williams, Brian Karrer +1

Large language models have increasingly been proposed as a powerful replacement for classical agent-based models (ABMs) to simulate social dynamics. By using LLMs as a proxy for hu…

cs.CL2022

Learning from data in the mixed adversarial non-adversarial case: Finding the helpers and ignoring the trolls

Da Ju, Jing Xu, Y-Lan Boureau +1

The promise of interaction between intelligent conversational agents and humans is that models can learn from such feedback in order to improve. Unfortunately, such exchanges in th…

cs.CL2020

Open-Domain Conversational Agents: Current Progress, Open Problems, and Future Directions

Stephen Roller, Y-Lan Boureau, Jason Weston +13

We present our view of what is necessary to build an engaging open-domain conversational agent: covering the qualities of such an agent, the pieces of the puzzle that have been bui…

cs.CL2025

Domain Regeneration: How well do LLMs match syntactic properties of text domains?

Da Ju, Hagen Blix, Adina Williams

Recent improvement in large language model performance have, in all likelihood, been accompanied by improvement in how well they can approximate the distribution of their training…