benchmark evaluation 1conflict resolution 1large language models 1specification analysis 1symmetry design 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
Tairan Wang, Liang Zhou, Zikang Zhan +1
The paper presents a symmetry‑based experimental framework for measuring how large language models resolve conflicts between competing specifications, and evaluates systematic pref…
cs.CL2024
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
Jianhao Yan, Pingchuan Yan, Yulong Chen +3
This study presents a comprehensive evaluation of GPT-4's translation capabilities compared to human translators of varying expertise levels. Through systematic human evaluation us…