activity
20132023
most citedDS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

33 citations · 85 across the 13 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.LG20227 cited

Coder Reviewer Reranking for Code Generation

Tianyi Zhang, Tao Yu, Tatsunori B. Hashimoto +4

Sampling diverse programs from a code language model and reranking with model likelihood is a popular method for code generation but it is prone to preferring degenerate solutions.…

cs.CV20221 cited

G^3: Geolocation via Guidebook Grounding

Grace Luo, Giscard Biamby, Trevor Darrell +2

We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks th…

cs.CL2022

AutoReply: Detecting Nonsense in Dialogue Introspectively with Discriminative Replies

Weiyan Shi, Emily Dinan, Adi Renduchintala +4

Existing approaches built separate classifiers to detect nonsense in dialogues. In this paper, we show that without external classifiers, dialogue models can detect errors in their…

cs.SE202233 cited

DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Yuhang Lai, Chengxi Li, Yiming Wang +7

We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as NumPy and Pandas. Compared to prior works, DS-1000…

cs.CL20221 cited

Inferring Rewards from Language in Context

Jessy Lin, Daniel Fried, Dan Klein +1

In classic instruction following, language like "I'd like the JetBlue flight" maps to actions (e.g., selecting that flight). However, language also conveys information about a user…