activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

Huichao Zhang, Liao Qu, Yiheng Liu +33

We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…

cs.CV2025

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

Jiahao Wang, Yufeng Yuan, Rujie Zheng +12

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current…

cs.CV2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

Xianfeng Tan, Yuhan Li, Wenxiang Shang +6

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-wo…

cs.CV2025

Compress Any Segment Anything Model (SAM)

Juntong Fan, Zhiwei Hao, Jianqiang Shen +4

Due to the excellent performance in yielding high-quality, zero-shot segmentation, Segment Anything Model (SAM) and its variants have been widely applied in diverse scenarios such…

cs.CV2025

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset

Zhuowei Chen, Bingchuan Li, Tianxiang Ma +8

Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructi…

cs.CV2024

Multi-Garment Customized Model Generation

Yichen Liu, Penghui Du, Yi Liu Quanwei Zhang

This paper introduces Multi-Garment Customized Model Generation, a unified framework based on Latent Diffusion Models (LDMs) aimed at addressing the unexplored task of synthesizing…