5 papers
Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation
Chen Li, Peng Zhang, Hanyu Zhou +5
Streaming autoregressive video models generate long videos chunk by chunk, using historical memory to maintain consistency. Existing methods typically expose subject and scene quer…
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
Yilei Hua, Beibei Jing, Ce Zheng +3
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods…
Achieving Text-based Person Retrieval with Any Granularity
Jialong Zuo, Hanyu Zhou, Dongyue Wu +5
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper introduces a new paradi…
Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets
Jialong Zuo, Haoyou Deng, Hanyu Zhou +10
The rapid evolution of text-to-image generation models has revolutionized visual content creation. While commercial products like Nano Banana Pro have garnered significant attentio…
Cross-video Identity Correlating for Person Re-identification Pre-training
Jialong Zuo, Ying Nie, Hanyu Zhou +5
Recent researches have proven that pre-training on large-scale person images extracted from internet videos is an effective way in learning better representations for person re-ide…