4 papers
WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model
Leyang Chen, Junyi Wu, Shaoqiu Zhang +1
Diffusion world models generate high-quality futures, but re- peated transformer evaluations make inference prohibitively slow. Existing caches reuse intermediate features, selecti…
GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction
Leyang Chen, Junyi Wu, Zhiteng Li +1
Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value (KV) cache that grows linearly with sequence length, creating a severe memory bottl…
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Junxian Li, Kai Liu, Leyang Chen +7
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
Kai Liu, Leyang Chen, Wenbo Li +5
Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. However, evaluations of unified multimodal models (UMMs) rem…