From the 1 of 6 linked papers with an AI index.
6 papers
Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang +26
The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Shuailei Ma, Jiaqi Liao, Xinyang Wang +24
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…
OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation
Yuheng Liu, Xin Lin, Xinke Li +9
Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthe…
LoST: Level of Semantics Tokenization for 3D Shapes
Niladri Shekhar Dutt, Zifan Shi, Paul Guerrero +4
Tokenization is a fundamental technique in the generative modeling of various modalities. In particular, it plays a critical role in autoregressive (AR) models, which have recently…
RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets
Isabella Liu, Zhan Xu, Wang Yifan +5
We present RigAnything, a novel autoregressive transformer-based model, which makes 3D assets rig-ready by probabilistically generating joints and skeleton topologies and assigning…
Learning Naturally Aggregated Appearance for Efficient 3D Editing
Ka Leong Cheng, Qiuyu Wang, Zifan Shi +5
Neural radiance fields, which represent a 3D scene as a color field and a density field, have demonstrated great progress in novel view synthesis yet are unfavorable for editing du…