6 papers
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
Zonghao Ying, Xiangfan Wu, Huiyu Wu +4
Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning,…
Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming
Yong Yang, Xing Zheng, Huiyu Wu +7
The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol (MCP) ecosystem and the language models themselves, h…
XiYan-SQL: A Novel Multi-Generator Framework For Text-to-SQL
Yifu Liu, Yin Zhu, Yingqi Gao +8
To leverage the advantages of LLM in addressing challenges in the Text-to-SQL task, we present XiYan-SQL, an innovative framework effectively generating and utilizing multiple SQL…
Evaluation Report on MCP Servers
Zhiling Luo, Xiaorong Shi, Xuanrui Lin +1
With the rise of LLMs, a large number of Model Context Protocol (MCP) services have emerged since the end of 2024. However, the effectiveness and efficiency of MCP servers have not…
A Preview of XiYan-SQL: A Multi-Generator Ensemble Framework for Text-to-SQL
Yingqi Gao, Yifu Liu, Xiaoxia Li +10
To tackle the challenges of large language model performance in natural language to SQL tasks, we introduce XiYan-SQL, an innovative framework that employs a multi-generator ensemb…
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
Pengfei Jing, Mengyun Tang, Xiaorong Shi +5
Evaluating Large Language Models (LLMs) is crucial for understanding their capabilities and limitations across various applications, including natural language processing and code…