3 papers
cs.AI2025
Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess Evaluation
Xidan Song, Weiqi Wang, Ruifeng Cao +1
The evaluation of Large Language Models (LLMs) in complex reasoning domains typically relies on performance alignment with ground-truth oracles. In the domain of chess, this standa…
cs.CL2025
Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement
Jing Ren, Weiqi Wang
Large language models (LLMs) like ChatGPT are increasingly used in academic writing, yet issues such as incorrect or fabricated references raise ethical concerns. Moreover, current…
cs.SE2025
Supporting Software Formal Verification with Large Language Models: An Experimental Study
Weiqi Wang, Marie Farrell, Lucas C. Cordeiro +1
Formal methods have been employed for requirements verification for a long time. However, it is difficult to automatically derive properties from natural language requirements. Spe…