activity
20242026
collaborators

6 papers

cs.CL2026

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Jiahao Ying, Boxian Ai, Wei Tang +2

Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improving agent performance on real-…

cs.AI2025

FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization

Shibo Hong, Jiahao Ying, Haiyuan Liang +4

Evaluating open-ended outputs of Multimodal Large Language Models has become a bottleneck as model capabilities, task diversity, and modality rapidly expand. Existing ``MLLM-as-a-J…

cs.LG2025

Beyond Benchmarks: Understanding Mixture-of-Experts Models through Internal Mechanisms

Jiahao Ying, Mingbao Lin, Qianru Sun +1

Mixture-of-Experts (MoE) architectures have emerged as a promising direction, offering efficiency and scalability by activating only a subset of parameters during inference. Howeve…

cs.CL2025

EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization

Yaoning Wang, Jiahao Ying, Yixin Cao +2

The rapid advancement of large language models (LLMs) and the development of increasingly large and diverse evaluation benchmarks have introduced substantial computational challeng…

cs.CL2025

Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric

Yixin Cao, Jiahao Ying, Yaoning Wang +3

Large Language Models (LLMs) have become indispensable across academia, industry, and daily applications, yet current evaluation methods struggle to keep pace with their rapid deve…

cs.CL2025

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Yixin Cao, Shibo Hong, Xinze Li +24

Large Language Models (LLMs) are advancing at an amazing speed and have become indispensable across academia, industry, and daily applications. To keep pace with the status quo, th…