6 papers
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
Yiming Pan, Chengwei Hu, Xuancheng Huang +6
Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a fundamental challenge: the gen…
WiS Platform: Enhancing Evaluation of LLM-Based Multi-Agent Systems Through Game-Based Analysis
Chengwei Hu, Jianhui Zheng, Yancheng He +7
Recent advancements in autonomous multi-agent systems (MAS) based on large language models (LLMs) have enhanced the application scenarios and improved the capability of LLMs to han…
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
Haibin Chen, Kangtao Lv, Chengwei Hu +8
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilit…
AIR: Complex Instruction Generation via Automatic Iterative Refinement
Wei Liu, Yancheng He, Hui Huang +5
With the development of large language models, their ability to follow simple instructions has significantly improved. However, adhering to complex instructions remains a major cha…
Mitigating Out-of-Entity Errors in Named Entity Recognition: A Sentence-Level Strategy
Guochao Jiang, Ziqin Luo, Chengwei Hu +2
Many previous models of named entity recognition (NER) suffer from the problem of Out-of-Entity (OOE), i.e., the tokens in the entity mentions of the test samples have not appeared…
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
Yancheng He, Shilong Li, Jiaheng Liu +15
New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive…