works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.SE2026

KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models

Fuyuan Xia, Qixin Zhang, Chenhao Ying +5

The paper introduces KQFuzz, a knowledge-guided fuzzing framework that uses large language models to generate and mutate test programs for quantum libraries, achieving higher cover…

cs.SE2026

AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

Pingchuan Ma, Zhaoyu Wang, Zimo Ji +5

Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, their autonomy poses significant…

cs.SE2026

SkillReducer: Optimizing LLM Agent Skills for Token Efficiency

Yudong Gao, Zongjie Li, Yuanyuan Yuan +3

LLM-based coding agents rely on \emph{skills}, pre-packaged instruction sets that extend agent capabilities, yet every token of skill content injected into the context window incur…

cs.CR2026

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails

Yuguang Zhou, Xunguang Wang, Pingchuan Ma +3

LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents. However, we reveal that the very reasoning and…

cs.AI2026

Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models

Xunguang Wang, Yuguang Zhou, Qingyue Wang +5

Large language models increasingly rely on explicit chain-of-thought reasoning to solve complex tasks, yet the safety of the reasoning process itself remains largely unaddressed. E…

cs.CY2026

WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making

Zongjie Li, Chaozheng Wang, Yuchong Xie +2

Large Language Models are increasingly being considered for deployment in safety-critical military applications. However, current benchmarks suffer from structural blindspots that…