Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
Chang Yang, Ruiyu Wang, Junzhe Jiang +9
Reasoning is the fundamental capability of large language models (LLMs). Due to the rapid progress of LLMs, there are two main issues of current benchmarks: i) these benchmarks can…
cs.AI2025
FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
Junzhe Jiang, Chang Yang, Aixin Cui +6
Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, an…