Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
Yingji Zhang, Yong Dai, André Freitas
Mathematical reasoning is a hallmark of human intelligence, requiring logical deduction, symbolic manipulation, and abstract thinking. Recent multimodal large language models (MLLM…
cs.CL2026
A2Eval: Agentic and Automated Evaluation for Embodied Brain
Shuai Zhang, Jiayu Hu, Zijie Chen +9
Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm…