works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

Xuan He, Cong Wei, Yuhao Cheng +19

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: b…

cs.CV2026

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

Jiaxing Li, Kai Zou, Cindy Zhou +7

The paper studies autoregressive video distillation, showing that aligning the student model’s mode coverage with the teacher’s distribution improves generation quality and diversi…

cs.CV2026

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

Dian Zheng, Manyuan Zhang, Hongyu Li +7

Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention from researchers. The main challeng…

cs.CV2026

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

Kai Zou, Ziqi Huang, Yuhao Dong +7

Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat…

cs.CV2026

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

Kai Zou, Hongbo Liu, Dian Zheng +3

In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with accurate layouts and high fidelity to te…

cs.CV2025

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Dian Zheng, Ziqi Huang, Hongbo Liu +9

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate t…