3 papers
cs.CR2026
MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills
Wenbo Guo, Wei Zeng, Chengwei Liu +5
AI coding agents such as Claude Code and Gemini CLI increasingly extend themselves with third-party skills: markdown packages bundling natural-language instructions, executable scr…
cs.AI2026
GameUIAgent: An LLM-Powered Framework for Automated Game UI Design with Structured Intermediate Representation
Wei Zeng, Fengwei An, Zhen Liu +1
Game UI design requires consistent visual assets across rarity tiers yet remains a predominantly manual process. We present GameUIAgent, an LLM-powered agentic framework that trans…
cs.SE2025
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
Ruizhong Qiu, Weiliang Will Zeng, James Ezick +2
The emergence of large language models (LLMs) has significantly pushed the frontiers of program synthesis. Advancement of LLM-based program synthesis calls for a thorough evaluatio…