activity
20232026
collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL2026

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

Zhouyuan Ma, Yutao Wu, Hanxun Huang +6

Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is kn…

cs.CL2026

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Jiahao Ying, Boxian Ai, Wei Tang +2

Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improving agent performance on real-…

cs.CL2025

EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization

Yaoning Wang, Jiahao Ying, Yixin Cao +2

The rapid advancement of large language models (LLMs) and the development of increasingly large and diverse evaluation benchmarks have introduced substantial computational challeng…

cs.CL2025

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

Jiahao Ying, Wei Tang, Yiran Zhao +3

This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic…

cs.CL2025

Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric

Yixin Cao, Jiahao Ying, Yaoning Wang +3

Large Language Models (LLMs) have become indispensable across academia, industry, and daily applications, yet current evaluation methods struggle to keep pace with their rapid deve…

cs.CL2025

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches

Yuhang Zhou, Xutian Chen, Yixin Cao +8

Recent progress in large language models (LLMs) has outpaced the development of effective evaluation methods. Traditional benchmarks rely on task-specific metrics and static datase…