4 papers
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Yixuan Lai, Tianjia Shao, Kun Zhou +3
GroundShot is a training-free, model-agnostic framework that improves visual consistency in multi-shot video generation by maintaining an online entity-level visual memory and sche…
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
Weijia Dou, Wenzhao Zheng, Weiliang Chen +3
Recent generative models can produce high-fidelity videos, yet they often exhibit 3D spatial geometric inconsistencies. Existing evaluation methods fail to accurately characterize…
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
Weijia Dou, Hui Li, Jiahao Cui +3
Streaming video generation models typically rely on temporal-centric memory, which organizes historical context as raw frames, chunk segments, or unclustered tokens. This organizat…
GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
Weijia Dou, Xu Zhang, Yi Bin +5
Recent attempts to transfer features from 2D Vision-Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields…