2 papers
cs.SE2026
ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair
Zhiyong Chen, Jialun Cao, Chang Xu +1
Model checking in TLA+ provides strong correctness guarantees, yet practitioners continue to face significant challenges in interpreting counterexamples, understanding large state-…
cs.LG2024
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
Jialun Cao, Zhiyong Chen, Jiarong Wu +2
Code generation benchmarks such as HumanEval are widely adopted to evaluate LLMs' capabilities. However, after consolidating the latest 24 benchmarks, we noticed three significant…