11 papers
Hit-and-Run Mixes as Fast as the Ball Walk
Ruizhe Zhang
Let be an isotropic convex body. We prove that the hit-and-run walk, started from any -warm distribution, reaches total-variation distance f…
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Junhao Chen, Mingjin Chen, Jingjia Mao +12
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Junhao Chen, Mingjin Chen, Henghaofan Zhang +10
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +15
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interac…
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
Zhantao Gong, Liaoyuan Fan, Qing Guo +3
In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet lar…
AlayaWorld: Long-Horizon and Playable Video World Generation
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +14
Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after dep…