collaborators

19 papers

cs.CV2026

Let RGB Be the Language of Vision

Timing Yang, Jinrui Yang, Xinlong Li +11

The paper proposes a unified vision framework that encodes all visual signals—including images, masks, and depth maps—as RGB images, turning diverse tasks into a common RGB-to-RGB…

cs.CV2026

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

Haoyu Chen, Kaichen Zhou, Hang Hua +11

The paper introduces MemoBench, a benchmark that tests video generation models' ability to remember and correctly update objects that disappear and later reappear in dynamically ch…

cs.CV2026

Large Language Models are Universal Reasoners for Visual Generation

Sucheng Ren, Chen Chen, Zhenbang Wang +5

Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM backbone handles both visual un…

cs.CV2026

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

Xingrui Wang, Jiang Liu, Chao Huang +7

Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks primarily evaluate general cross-mo…

cs.CV2026

Captain Safari: A World Engine with Pose-Aligned 3D Memory

Yu-Cheng Chou, Xingrui Wang, Yitong Li +5

World engines aim to synthesize long, 3D-consistent videos that support interactive exploration of a scene under user-controlled camera motion. However, existing systems struggle u…

cs.CV2026

VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents

Feng Wang, Yichun Shi, Ceyuan Yang +4

This work presents VTok, a unified video tokenization framework that can be used for both generation and understanding tasks. Unlike the leading vision-language systems that tokeni…