activity
20242026
most citedSCBench: A Sports Commentary Benchmark for Video LLMs

2 citations · 2 across the 11 of their papers we have counts for

collaborators
Showing 2026Show all

7 papers · 1 filter

cs.AI2026

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +15

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the pr…

cs.SD2026

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Junhao Chen, Mingjin Chen, Jingjia Mao +12

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…

cs.CV2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

Junhao Chen, Mingjin Chen, Henghaofan Zhang +9

Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…

cs.AI2026

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +15

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interac…

cs.CV2026

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +14

Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after dep…

cs.RO2026

ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control

Haosen Wang, Zhenyang Li, Yinqiang Zhang +9

Zero-shot object navigation has advanced rapidly with open-vocabulary detectors, image--text models, and language-guided exploration. However, even after current methods detect a p…