9 papers
SceneActBench: Can Agents Act on the 3D Scenes They See?
Yifei Zhao, Xiangxin Zhou, Wenhao Yang +11
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operat…
Structured Scaling of AI Discovery Across Diverse Scientific Domains
Haotian Ye, Haowei Lin, Jingyi Tang +30
Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions. Language models can increasingly participate in these loops, but simply gen…
Can Language Models Discover Scaling Laws?
Haowei Lin, Haotian Ye, Wenzheng Feng +8
Discovering scaling laws for predicting model performance at scale is a fundamental and open-ended challenge, mostly reliant on slow, case specific human experimentation. To invest…
A Neural Symbolic Model for Space Physics
Jie Ying, Haowei Lin, Chao Yue +7
In this study, we unveil a new AI model, termed PhyE2E, to discover physical formulas through symbolic regression. PhyE2E simplifies symbolic regression by decomposing it into sub-…
Inference-time Scaling of Diffusion Models through Classical Search
Xiangcheng Zhang, Haowei Lin, Haotian Ye +4
Classical search algorithms have long underpinned modern artificial intelligence. In this work, we tackle the challenge of inference-time control in diffusion models -- adapting ge…
MCU: An Evaluation Framework for Open-Ended Game Agents
Xinyue Zheng, Haowei Lin, Kaichen He +3
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains diffic…