6 papers
SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
Dexu Yu, Youhua Li, Zhaoyang Guan +12
Agent skills have become a practical way to extend large language model agents, but the growing skill ecosystem still lacks a reliable way to judge whether a skill is worth deployi…
Symphony-Coord: Adaptive Routing for Multi-Agent LLM Systems
Zhaoyang Guan, Huixi Cao, Ming Zhong +6
Multi-agent large language model systems can tackle complex multi-step tasks by decomposing work and coordinating specialized behaviors. However, current coordination mechanisms ty…
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
Tianyi Jiang, Arctanx An, Hengyi Feng +12
Human problem-solving is never the repetition of a single mindset, by which we mean a distinct mode of cognitive processing. When tackling a specific task, we do not rely on a sing…
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
Yutong Bian, Xianhao Lin, Yupeng Xie +11
Large Language Models (LLMs) and code agents in software development are rapidly evolving from generating isolated code snippets to producing full-fledged software applications wit…
Teach Me How to Denoise: A Universal Framework for Denoising Multi-modal Recommender Systems via Guided Calibration
Hongji Li, Hanwen Du, Youhua Li +5
The surge in multimedia content has led to the development of Multi-Modal Recommender Systems (MMRecs), which use diverse modalities such as text, images, videos, and audio for mor…
Uncertainty-aware Knowledge Tracing
Weihua Cheng, Hanwen Du, Chunxiao Li +4
Knowledge Tracing (KT) is crucial in education assessment, which focuses on depicting students' learning states and assessing students' mastery of subjects. With the rise of modern…