3 papers
cs.AI2025
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
Shima Imani, Seungwhan Moon, Adel Ahmadyan +3
We introduce, a large-scale synthetic benchmark of 15,045 university-level physics problems (90/10% train/test split). Each problem is fully parameterized, supporting an effectivel…
cs.AI2025
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
Shima Imani, Seungwhan Moon, Adel Ahmadyan +3
Evaluating vision-language models (VLMs) in scientific domains like mathematics and physics poses unique challenges that go far beyond predicting final answers. These domains deman…
cs.AI2025
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
Eun Chang, Zhuangqun Huang, Yiwei Liao +19
We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like sm…