activity
20202024
most citedCan Large Language Models Be an Alternative to Human Evaluations?

36 citations · 50 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2024

Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Guan-Ting Lin, Cheng-Han Chiang, Hung-yi Lee

In spoken dialogue, even if two current turns are the same sentence, their responses might still differ when they are spoken in different styles. The spoken styles, containing para…

cs.CL2024

Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations

Cheng-Han Chiang, Hung-yi Lee

Long-form generations from large language models (LLMs) contain a mix of factual and non-factual claims, making evaluating factuality difficult. Prior works evaluate the factuality…

cs.CL20241 cited

Over-Reasoning and Redundant Calculation of Large Language Models

Cheng-Han Chiang, Hung-yi Lee

Large language models (LLMs) can solve problems step-by-step. While this chain-of-thought (CoT) reasoning boosts LLMs' performance, it is unclear if LLMs \textit{know} when to use…

cs.CL20234 cited

A Closer Look into Automatic Evaluation Using Large Language Models

Cheng-Han Chiang, Hung-yi Lee

Using large language models (LLMs) to evaluate text quality has recently gained popularity. Some prior works explore the idea of using LLMs for evaluation, while they differ in som…

cs.CL2023

Revealing the Blind Spot of Sentence Encoder Evaluation by HEROS

Cheng-Han Chiang, Yung-Sung Chuang, James Glass +1

Existing sentence textual similarity benchmark datasets only use a single number to summarize how similar the sentence encoder's decision is to humans'. However, it is unclear what…

cs.CL202336 cited

Can Large Language Models Be an Alternative to Human Evaluations?

Cheng-Han Chiang, Hung-yi Lee

Human evaluation is indispensable and inevitable for assessing the quality of texts generated by machine learning models or written by humans. However, human evaluation is very dif…