4 papers · 1 filter
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
Haibin Chen, Kangtao Lv, Chengwei Hu +8
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilit…
AIR: Complex Instruction Generation via Automatic Iterative Refinement
Wei Liu, Yancheng He, Hui Huang +5
With the development of large language models, their ability to follow simple instructions has significantly improved. However, adhering to complex instructions remains a major cha…
Mitigating Out-of-Entity Errors in Named Entity Recognition: A Sentence-Level Strategy
Guochao Jiang, Ziqin Luo, Chengwei Hu +2
Many previous models of named entity recognition (NER) suffer from the problem of Out-of-Entity (OOE), i.e., the tokens in the entity mentions of the test samples have not appeared…
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
Yancheng He, Shilong Li, Jiaheng Liu +15
New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive…