most citedImproving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback

35 citations · 37 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL20241 cited

Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning

Danna Zheng, Mirella Lapata, Jeff Z. Pan

We present Archer, a challenging bilingual text-to-SQL dataset specific to complex reasoning, including arithmetic, commonsense and hypothetical reasoning. It contains 1,042 Englis…

cs.CL2024

Improving Generalization in Semantic Parsing by Increasing Natural Language Variation

Irina Saparina, Mirella Lapata

Text-to-SQL semantic parsing has made significant progress in recent years, with various models demonstrating impressive performance on the challenging Spider benchmark. However, i…

cs.CL2024

TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness

Danna Zheng, Danyang Liu, Mirella Lapata +1

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, prompting a surge in their practical applications. However, concerns have arisen rega…

cs.CL2023

Visual Storytelling with Question-Answer Plans

Danyang Liu, Mirella Lapata, Frank Keller

Visual storytelling aims to generate compelling narratives from image sequences. Existing models often focus on enhancing the representation of the image sequence, e.g., with exter…

cs.CL2023

Attributable and Scalable Opinion Summarization

Tom Hosking, Hao Tang, Mirella Lapata

We propose a method for unsupervised opinion summarization that encodes sentences from customer reviews into a hierarchical discrete latent space, then identifies common opinions b…

cs.CL202335 cited

Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback

Yao Fu, Hao Peng, Tushar Khot +1

We study whether multiple large language models (LLMs) can autonomously improve each other in a negotiation game by playing, reflecting, and criticizing. We are interested in this…