2 papers
cs.SE2025
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
Honghua Dong, Jiacheng Yang, Xun Deng +4
Type inference for dynamic languages like Python is a persistent challenge in software engineering. While large language models (LLMs) have shown promise in code understanding, the…
cs.SE2025
VerifyThisBench: Generating Code, Specifications, and Proofs All at Once
Xun Deng, Sicheng Zhong, Barış Bayazıt +3
Large language models (LLMs) have demonstrated remarkable progress in code generation, but many existing benchmarks are approaching saturation and offer little guarantee on the tru…