most citedMultiple-Choice Questions are Efficient and Robust LLM Evaluators

3 citations · 5 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL20243 cited

Multiple-Choice Questions are Efficient and Robust LLM Evaluators

Ziyin Zhang, Zhaokun Jiang, Lizhen Xu +2

We present GSM-MC, a multiple-choice (MC) dataset constructed by collecting answers and incorrect predictions on GSM8K from 60 open-source models. Through extensive experiments, we…

cs.CL2024

Improving Open-Ended Text Generation via Adaptive Decoding

Wenhong Zhu, Hongkun Hao, Zhiwei He +2

Current language models decode text token by token according to probabilistic distribution, and determining the appropriate candidates for the next token is crucial to ensure gener…

cs.CL2024

Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models

Zhiwei He, Binglin Zhou, Hongkun Hao +5

Text watermarking technology aims to tag and identify content produced by large language models (LLMs) to prevent misuse. In this study, we introduce the concept of cross-lingual c…

cs.CL20232 cited

Boosting Large Language Model for Speech Synthesis: An Empirical Study

Hongkun Hao, Long Zhou, Shujie Liu +4

Large language models (LLMs) have made significant advancements in natural language processing and are concurrently extending the language ability to other modalities, such as spee…

cs.CL2023

CLEAN-EVAL: Clean Evaluation on Contaminated Large Language Models

Wenhong Zhu, Hongkun Hao, Zhiwei He +6

We are currently in an era of fierce competition among various large language models (LLMs) continuously pushing the boundaries of benchmark performance. However, genuinely assessi…

cs.CL2023

Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text Generation

Wenhong Zhu, Hongkun Hao, Rui Wang

The decoding algorithm is critical for open-ended text generation, transforming latent representations into coherent and meaningful outputs. This paper investigates the self-reinfo…