2 papers
cs.AI2026
Benchmarking Small Language Models and Small Reasoning Language Models on System Log Severity Classification
Yahya Masri, Emily Ma, Zifu Wang +2
System logs are crucial for monitoring and diagnosing modern computing infrastructure, but their scale and complexity require reliable and efficient automated interpretation. Since…
cs.CG2025
SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
Peijie Wang, Chao Yang, Zhong-Zhi Li +6
Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing mul…