189 citations · 253 across the 21 of their papers we have counts for
6 papers · 1 filter
PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion
Yiduo Guo, Zekai Zhang, Yaobo Liang +2
Recent evaluations of Large Language Models (LLMs) have centered around testing their zero-shot/few-shot capabilities for basic natural language tasks and their ability to translat…
EIPE-text: Evaluation-Guided Iterative Plan Extraction for Long-Form Narrative Text Generation
Wang You, Wenshan Wu, Yaobo Liang +8
Plan-and-Write is a common hierarchical approach in long-form narrative text generation, which first creates a plan to guide the narrative writing. Following this approach, several…
GameEval: Evaluating LLMs on Conversational Games
Dan Qiao, Chenfei Wu, Yaobo Liang +2
The rapid advancements in large language models (LLMs) have presented challenges in evaluating those models. Existing evaluation methods are either reference-based or preference ba…
Analyzing and Reducing the Performance Gap in Cross-Lingual Transfer with Fine-tuning Slow and Fast
Yiduo Guo, Yaobo Liang, Dongyan Zhao +2
Existing research has shown that a multilingual pre-trained language model fine-tuned with one (source) language also performs well on downstream tasks for non-source languages, ev…
Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models
Zhihong Shao, Yeyun Gong, Yelong Shen +3
Large language models can perform various reasoning tasks by using chain-of-thought prompting, which guides them to find answers through step-by-step demonstrations. However, the q…
Improving Task Generalization via Unified Schema Prompt
Wanjun Zhong, Yifan Gao, Ning Ding +5
Task generalization has been a long standing challenge in Natural Language Processing (NLP). Recent research attempts to improve the task generalization ability of pre-trained lang…