2 papers
cs.CL2026
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
Yihong Dong, Jianha Xiao, Xue Jiang +7
The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack systematic evaluation based on comput…
cs.SE2025
LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding
Jia Li, Xuyuan Guo, Lei Li +7
Current advanced long-context language models offer great potential for real-world software engineering applications. However, progress in this critical domain remains hampered by…