3 papers
cs.CV2026
AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer
Junqiu Yu, Pandeng Li, Yikai Wang +11
Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoisi…
cs.CV2025
Online Dense Point Tracking with Streaming Memory
Qiaole Dong, Yanwei Fu
Dense point tracking is a challenging task requiring the continuous tracking of every point in the initial frame throughout a substantial portion of a video, even in the presence o…
cs.CV2025
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks
Ming Xie, Chenjie Cao, Yunuo Cai +3
In this paper, we present a novel Left-Prompt-Guided (LPG) paradigm to address a diverse range of reference-based vision tasks. Inspired by the human creative process, we reformula…