2 citations · 4 across the 5 of their papers we have counts for
4 papers · 1 filter
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Yunhan Wang, Jiaan Wang, Lianzhe Huang +2
Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing benchmarks such as BrowseComp rely…
Instruction Position Matters in Sequence Generation with Large Language Models
Yijin Liu, Xianfeng Zeng, Fandong Meng +1
Large language models (LLMs) are capable of performing conditional sequence generation tasks, such as translation or summarization, through instruction fine-tuning. The fine-tuning…
Towards Multiple References Era -- Addressing Data Leakage and Limited Reference Diversity in NLG Evaluation
Xianfeng Zeng, Yijin Liu, Fandong Meng +1
N-gram matching-based evaluation metrics, such as BLEU and chrF, are widely utilized across a range of natural language generation (NLG) tasks. However, recent studies have reveale…
WeChat Neural Machine Translation Systems for WMT21
Xianfeng Zeng, Yijin Liu, Ernan Li +5
This paper introduces WeChat AI's participation in WMT 2021 shared news translation task on English->Chinese, English->Japanese, Japanese->English and English->German. Our systems…