2 papers
cs.AI2026
Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities
Xu Zhang, Xudong Gong, Jiacheng Qin +5
Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation, limiting targeted model im…
cs.AI2025
Pay More Attention to the Robustness of Prompt for Instruction Data Mining
Qiang Wang, Dawei Feng, Xu Zhang +4
Instruction tuning has emerged as a paramount method for tailoring the behaviors of LLMs. Recent work has unveiled the potential for LLMs to achieve high performance through fine-t…