works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

Xiangbo Gao, Siyuan Yang, Ping He +12

Visko Orbis 1.0 is a live model that generates long videos in real time, letting users change prompts on the fly while preserving subject, scene, and style consistency across hour‑…

cs.CV2026

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Shijie Li, Yilin Gao, Siyuan Yang +7

Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuanc…

cs.CV2026

CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

Fangzhou Lin, Peiran Li, Lingyu Xu +12

Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture th…

cs.CL2026

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

Hao Liu, Siyuan Yang, Qinglei Hu +1

Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth orbits grow crowded and conte…

cs.AI2026

CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning

Fangzhou Lin, Shuo Xing, Peiran Li +6

Parallel reasoning, where a generator samples many candidate solutions and an aggregator selects the best, is one of the most effective forms of test-time scaling in large language…

cs.CV2026

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Xiangbo Gao, Sicong Jiang, Bangya Liu +12

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional…