5 papers
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Qiucheng Yu, Ruijie Xu, Mingang Chen +2
The paper introduces TSHA, a large benchmark of real-world indoor safety hazard assessment questions for evaluating vision‑language models, and shows that training on this data imp…
MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction
Ruijie Xu, Xinnan Zhu, Jiayu Ying +3
Text-driven 3D scene generation is a promising technique for digital content creation, embodied AI simulation, and interactive design, yet practical workflows often require refinin…
Rethinking Data Mixing from the Perspective of Large Language Models
Yuanjian Xu, Tianze Sun, Changwei Xu +7
Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly reduce generalization. Although…
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
Yucheng Zeng, Shupeng Li, Daxiang Dong +11
Progress in software-engineering agents is increasingly constrained by the scarcity of executable, scalable, and realistic data for training and evaluation. This scarcity stems fro…
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Zhen Huang, Zengzhi Wang, Shijie Xia +25
The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showc…