most citedPersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

physics.soc-ph2025

FlockVote: LLM-Empowered Agent-Based Modeling for Simulating U.S. Presidential Elections

Lingfeng Zhou, Yi Xu, Zhenyu Wang +1

Modeling complex human behavior, such as voter decisions in national elections, is a long-standing challenge for computational social science. Traditional agent-based models (ABMs)…

cs.DC2025

FlowMesh: A Service Fabric for Composable LLM Workflows

Junyi Shen, Noppanat Wadlom, Lingfeng Zhou +4

AI deployment increasingly resembles a pipeline of data transformation, fine-tuning, and agent interactions rather than a monolithic LLM job; recent examples include RLHF/RLAIF tra…

cs.CL2025

MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding

Mohan Jiang, Jin Gao, Jiahao Zhan +1

As multimodal large language models (MLLMs) grow increasingly capable, fixed benchmarks are gradually losing their effectiveness in evaluating high-level scientific understanding.…

cs.AI2025

DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery

Keyu Li, Mohan Jiang, Dayuan Fu +4

The rapid advancement of large language models has fundamentally shifted the bottleneck in AI development from computational power to data availability-with countless valuable data…

cs.CL20251 cited

PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?

Lingfeng Zhou, Jialing Zhang, Jin Gao +2

Current role-play studies often rely on unvalidated LLM-as-a-judge paradigms, which may fail to reflect how humans perceive role fidelity. A key prerequisite for human-aligned eval…