4 papers
ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
Zhuohan Wang, Ziwei Zhu, Ziniu Li +8
Formulating optimization problems for industrial applications demands significant manual effort and domain expertise. While Large Language Models (LLMs) show promise in automating…
Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing
Luan Vejsiu, Qianyu Zheng, Haoxuan Chen +1
Despite ASR technology being full-scale adopted by industry and for large portions of the population, ASR systems often have errors that require editors to post-edit text quality.…
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
Qiyue Yin, Pei Xu, Qiaozhe Li +15
Recent breakthroughs in Large Language Models (LLMs) have led to a qualitative leap in artificial intelligence' s performance on reasoning tasks, particularly demonstrating remarka…
Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation
Zebin Wang, Yi Han, Ethan X. Fang +2
We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs h…