most citedA Survey on Large Language Model Benchmarks

2 citations · 2 across the 4 of their papers we have counts for

collaborators

8 papers

cs.SE2026

LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

Dong Xu, Jialun Cao, Guozhao Mo +9

Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Although large language models (LL…

cs.CL2026

All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG

Dan Wang, Guozhao Mo, Yafei Shi +9

Multilingual Retrieval-Augmented Generation (mRAG) leverages cross-lingual evidence to ground Large Language Models (LLMs) in global knowledge. However, we show that current mRAG s…

cs.AI2026

DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation

Hao Zheng, Guozhao Mo, Xinru Yan +7

Presentation generation requires deep content research, coherent visual design, and iterative refinement based on observation. However, existing presentation agents often rely on p…

cs.CL2025

AI-Salesman: Towards Reliable Large Language Model Driven Telemarketing

Qingyu Zhang, Chunlei Xin, Xuanang Chen +7

Goal-driven persuasive dialogue, exemplified by applications like telemarketing, requires sophisticated multi-turn planning and strict factual faithfulness, which remains a signifi…

cs.CL2025

CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection

Yihan Chen, Jiawei Chen, Guozhao Mo +4

The growing integration of large language models (LLMs) into the peer review process presents potential risks to the fairness and reliability of scholarly evaluation. While LLMs of…

cs.CL20252 cited

A Survey on Large Language Model Benchmarks

Shiwen Ni, Guhong Chen, Shuaimin Li +11

In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in incre…