Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
Shengyu Guo, Tongrui Ye, Jianbo Zhang +3
Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence.…
cs.AI2026
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
Xiaoxiao Wang, Chunxiao Li, Junying Wang +6
As comprehensive large model evaluation becomes prohibitively expensive, predicting model performance from limited observations has become essential. However, existing statistical…