From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
Lingkai Kong, Zijian Wu, Yuzhe Gu +10
The paper introduces AdvancedMathBench, a benchmark suite for evaluating large language models on generating and verifying advanced mathematical proofs, and provides an automatic v…
cs.CV2026
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Peiming Li, Yifan Wang, Xiaotian Zhang +4
Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduct…