2 papers
cs.CR2026
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
Yanlin Wang, Ziyao Zhang, Chong Wang +5
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, but their proficiency in producing secure code remains a critical, under-explored area. E…
cs.SE2025
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
Yanlin Wang, Xinyi Xu, Jiachi Chen +3
The rise of large language models (LLMs) has sparked a surge of interest in agents, leading to the rapid growth of agent frameworks. Agent frameworks are software toolkits and libr…