collaborators

11 papers

cs.AI2026

ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration

Yang Liu, Shiwei Hou, Xiyuan Chen +13

EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM…

cs.CR2026

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery

Muxi Lyu, Karen Shieh, Yiwei Hou +3

Cross-Site Scripting (XSS) remains one of the most prevalent and damaging classes of web vulnerabilities. LLM-based coding agents offer a promising approach to XSS discovery by com…

cs.CR2026

Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases

Yiwei Hou, Hao Wang, Muxi Lyu +6

Memory safety vulnerabilities remain a significant threat even for projects with extensive fuzzing and manual auditing. Recent results suggest that large language models hold great…

cs.HC2026

Turning Intent into Specifications: A Benchmark and an Interactive User-Assistant Agent

Hao Wang, Ligong Han, Kai Xu +1

Today's agents are highly effective at implementing well-scoped software design plans, but user intent is often vague and admits multiple equally valid solutions. In this paper, we…

cond-mat.mtrl-sci2026

Rongzai agent: A Large Language Model-Based Autonomous Assistant for Rietveld Refinement of Neutron Diffraction Data

Qingmeng Li, Hao Wang, Dongbo Xiong +11

Neutron diffraction (ND) is an indispensable technique for determining atomic positions (especially light elements) and thus serves as a critical probe for revealing microscopic st…

cs.AI2026

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Hao Wang, Hanchen Li, Qiuyang Mang +3

Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a s…