3 papers
cs.CV2026
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
Yilei Hua, Beibei Jing, Ce Zheng +3
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods…
cs.CV2026
Achieving Text-based Person Retrieval with Any Granularity
Jialong Zuo, Hanyu Zhou, Dongyue Wu +5
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper introduces a new paradi…
cs.CV2025
Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets
Jialong Zuo, Haoyou Deng, Hanyu Zhou +10
The rapid evolution of text-to-image generation models has revolutionized visual content creation. While commercial products like Nano Banana Pro have garnered significant attentio…