13 papers
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
Bohan Lyu, Yucheng Yang, Siqiao Huang +25
Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities i…
OneRank: Unified Transformer-Native Ranking Architecture for Multi-Task Recommendation
Jiakai Tang, Sunhao Dai, Kun Wang +8
Multi-task learning (MTL) is essential in recommender systems to enable complementary learning among diverse user feedback. While modern industrial practices have shifted from DNNs…
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
Jiahao Wu, Ning Lu, Shengcai Liu +6
Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance perfor…
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
ShiYing Huang, Liang Lin, Yuer Li +6
In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tensi…
Mem-W: Latent Memory-Native GUI Agents
Guibin Zhang, Yaohui Ling, Fanci Meng +2
GUI agents are beginning to operate the web, mobile, and desktop as interactive worlds, where successful control depends on carrying forward visual, procedural, and task-level evid…
MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents
Zhengwei Xie, Zhisheng Chen, Ziyan Weng +7
Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to tran…