4 papers
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
Yalun Wu, Haotian Liu, Zhoujun Li +1
As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably re…
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
Junying Wang, Zicheng Zhang, Ye Shen +8
High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…
The Ever-Evolving Science Exam
Junying Wang, Zicheng Zhang, Yijin Guo +9
As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progr…
Affordance Benchmark for MLLMs
Junying Wang, Wenzhe Li, Yalun Wu +6
Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong…