22 papers
UniCode: Augmenting Evaluation for Code Reasoning
Xinyue Zheng, Haowei Lin, Shaofei Cai +3
The paper presents UniCode, a generative evaluation framework that augments seed coding problems and automatically generates tests to more rigorously assess large language models'…
Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
Kaichen He, Zihao Wang, Muyao Li +2
The paradigm of agentic AI is shifting from engineered complex workflows to post-training native models. However, existing agents are typically confined to static, predefined actio…
Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
Zhou Ziheng, Huacong Tang, Jinyuan Zhang +9
Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence, yet evaluating this capacit…
LoopNav: Benchmarking Spatial Consistency in World Models
Kewei Lian, Shaofei Cai, Yitao Liang +1
The ability to simulate the world in a spatially consistent manner is a crucial requirement for effective world models. Such a model enables high-quality visual generation, and als…
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
Guangyu Zhao, Kewei Lian, Haoxuan Ru +8
Goal-conditioned policies enable decision-making models to execute diverse behaviors based on specified goals, yet their downstream performance is often highly sensitive to the cho…
Can Language Models Discover Scaling Laws?
Haowei Lin, Haotian Ye, Wenzheng Feng +8
Discovering scaling laws for predicting model performance at scale is a fundamental and open-ended challenge, mostly reliant on slow, case specific human experimentation. To invest…