3 papers
cs.SE2026
TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
Shuangjie Yao, Hao Wang, Koushik Sen +3
Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on t…
cs.AI2026
Cost-Efficient Theorem Proving via Agent Orchestration in Program Verification
Shuangjie Yao, Nikolaus Holzer, Mark Paul Santolucito +3
Program verification establishes software correctness through machine-checkable proofs constructed in theorem provers. It's a guarantee especially valuable for code generated by la…
cs.CR2026
Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning
Zihan Zhang, Shuangjie Yao, Zesen Liu +8
Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity betwee…