From the 1 of 23 linked papers with an AI index.
23 papers
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Ruiming Liang, Yi Zhong, Yizhen Yuan +6
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinfo…
ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
Dongxiu Liu, Haoyi Niu, Peng Cheng +5
The paper presents ODEWorld, a continuous-time latent world model that learns a physical-time flow using ODEs to predict future states at arbitrary temporal resolutions, improving…
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Miracle Kang, Lights Shi, Lucy Liang +10
Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions prim…
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
Rongxu Cui, Zongzheng Zhang, Jingrui Pang +11
Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address th…
World Value Models for Robotic Manipulation
Zhihao Wang, Jianxiong Li, Yu Cui +4
Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. Mathematically, accurate value estimation demands deep temporal…
Horizon Adaptive Offline Policy Learning via Value Stitching
Kexin Zheng, Xianyuan Zhan, Xintao Yan
Learning accurate value functions plays a decisive role for reinforcement learning (RL) agents to solve long-horizon, complex tasks. Conventional temporal-difference (TD) learning…