31 citations · 74 across the 63 of their papers we have counts for
11 papers · 1 filter
SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior
Zhiyu Chen, Zihan Guo, Bo Huang +4
Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organ…
Fints: Efficient Inference-Time Personalization for LLMs with Fine-Grained Instance-Tailored Steering
Kounianhua Du, Jianxing Liu, Kangning Zhang +6
The rapid evolution of large language models (LLMs) has intensified the demand for effective personalization techniques that can adapt model behavior to individual user preferences…
APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
Jiarui Qin, Yunjia Xi, Junjie Huang +6
With the rapid development of LLM-based agents, there is a growing trend to incorporate agent-specific data into the pre-training stage of LLMs, aiming to better align LLMs with re…
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
Yuanyi Song, Heyuan Huang, Qiqiang Lin +9
The rapid advancement of multimodal large language models has enabled agents to operate mobile devices by directly interacting with graphical user interfaces, opening new possibili…
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
Lingyue Fu, Xin Ding, Linyue Pan +9
Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent's capabi…
Superplatforms Have to Attack AI Agents
Jianghao Lin, Jiachen Zhu, Zheli Zhou +4
Over the past decades, superplatforms, digital companies that integrate a vast range of third-party services and applications into a single, unified ecosystem, have built their for…