From the 1 of 9 linked papers with an AI index.
9 papers
SceneActBench: Can Agents Act on the 3D Scenes They See?
Yifei Zhao, Xiangxin Zhou, Wenhao Yang +11
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operat…
UniCode: Augmenting Evaluation for Code Reasoning
Xinyue Zheng, Haowei Lin, Shaofei Cai +3
The paper presents UniCode, a generative evaluation framework that augments seed coding problems and automatically generates tests to more rigorously assess large language models'…
Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
Zhou Ziheng, Huacong Tang, Jinyuan Zhang +9
Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence, yet evaluating this capacit…
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
Guangyu Zhao, Kewei Lian, Haoxuan Ru +8
Goal-conditioned policies enable decision-making models to execute diverse behaviors based on specified goals, yet their downstream performance is often highly sensitive to the cho…
Unified Cross-Scale 3D Generation and Understanding via Autoregressive Modeling
Shuqi Lu, Haowei Lin, Lin Yao +6
3D structure modeling is essential across scales, enabling applications from fluid simulation and 3D reconstruction to protein folding and molecular docking. Yet, despite shared 3D…
MCU: An Evaluation Framework for Open-Ended Game Agents
Xinyue Zheng, Haowei Lin, Kaichen He +3
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains diffic…