activity
20242026
collaborators

6 papers

cs.CV2026

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

Xiaojie Xu, Zhengyuan Lin, Runyi Li +3

Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models,…

cs.AI2026

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +15

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the pr…

cs.AI2026

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +15

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interac…

cs.CV2026

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +14

Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after dep…

cs.CV2025

Self-Evaluation Unlocks Any-Step Text-to-Image Generation

Xin Yu, Xiaojuan Qi, Zhengqi Li +6

We introduce the Self-Evaluating Model (Self-E), a novel, from-scratch training approach for text-to-image generation that supports any-step inference. Self-E learns from data simi…

cs.CV2024

Removing Distributional Discrepancies in Captions Improves Image-Text Alignment

Yuheng Li, Haotian Liu, Mu Cai +5

In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language mo…