Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
Zhenyu Li, Kehai Chen, Yunfei Long +5
Large Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks…
cs.CL2025
Generator-Assistant Stepwise Rollback Framework for Large Language Model Agent
Xingzuo Li, Kehai Chen, Yunfei Long +3
Large language model (LLM) agents typically adopt a step-by-step reasoning framework, in which they interleave the processes of thinking and acting to accomplish the given task. Ho…
cs.CL2024
Mitigating the Bias of Large Language Model Evaluation
Hongli Zhou, Hui Huang, Yunfei Long +5
Recently, there has been a trend of evaluating the Large Language Model (LLM) quality in the flavor of LLM-as-a-Judge, namely leveraging another LLM to evaluate the current output…