Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
Yang Liu, Shiwei Hou, Xiyuan Chen +13
EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM…
cs.AI2026
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Hao Wang, Hanchen Li, Qiuyang Mang +3
Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a s…