1 citations · 1 across the 5 of their papers we have counts for
7 papers
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
OMG: Omni-Modal Motion Generation for Generalist Humanoid Control
Siqiao Huang, Kun-Ying Lee, Dongming Qiao +5
Humanoid whole-body control has made significant progress in recent years, yet existing approaches remain limited to few-skill policies with heavy reward engineering, or motion tra…
Nano World Models: A Minimalist Implementation of Future Video Prediction
Siqiao Huang, Partha Kaushik, Michael Chen +5
World models have become a central paradigm for learning predictive simulators that support generation, planning, and decision-making. Yet, despite rapid progress in industry-scale…
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
Bohan Lyu, Yucheng Yang, Siqiao Huang +25
Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities i…
Vid2World: Crafting Video Diffusion Models to Interactive World Models
Siqiao Huang, Jialong Wu, Qixing Zhou +2
World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. How…
SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors
Bohan Lyu, Siqiao Huang, Zichen Liang
Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks, su…