1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2025
FreeInsert: Personalized Object Insertion with Geometric and Style Control
Yuhong Zhang, Han Wang, Yiwen Wang +2
Text-to-image diffusion models have made significant progress in image generation, allowing for effortless customized generation. However, existing image editing methods still face…
cs.CV2024★ 1 cited
Elysium: Exploring Object-level Perception in Videos via MLLM
Han Wang, Yanjie Wang, Yongjie Ye +2
Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking…
cs.CV2024
GloTSFormer: Global Video Text Spotting Transformer
Han Wang, Yanjie Wang, Yang Li +1
Video Text Spotting (VTS) is a fundamental visual task that aims to predict the trajectories and content of texts in a video. Previous works usually conduct local associations and…