8 papers
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
Rongxu Cui, Zongzheng Zhang, Jingrui Pang +11
Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address th…
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning
Wanshi Xu, Haokun Zhao, Haidong Yuan +2
Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxil…
Controllable Spoken Dialogue Generation: An LLM-Driven Grading System for K-12 Non-Native English Learners
Haidong Yuan, Haokun Zhao, Wanshi Xu +4
Large language models (LLMs) often fail to meet the pedagogical needs of K-12 English learners in non-native contexts due to a proficiency mismatch. To address this widespread chal…
Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning
Haokun Zhao, Wanshi Xu, Haidong Yuan +3
Geometric reasoning inherently requires "thinking with constructions" -- the dynamic manipulation of visual aids to bridge the gap between problem conditions and solutions. However…
Medical Test-free Disease Detection Based on Big Data
Haokun Zhao, Yingzhe Bai, Qingyang Xu +3
Accurate disease detection is of paramount importance for effective medical treatment and patient care. However, the process of disease detection is often associated with extensive…
TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis
Haokun Zhao, Xiang Zhang, Jiaqi Wei +4
Time series forecasting is central to decision-making in domains as diverse as energy, finance, climate, and public health. In practice, forecasters face thousands of short, noisy…