3 papers
cs.CV2025
Mean of Means: Human Localization with Calibration-free and Unconstrained Camera Settings (extended version)
Tianyi Zhang, Wengyu Zhang, Xulu Zhang +4
Accurate human localization is crucial for various applications, especially in the Metaverse era. Existing high precision solutions rely on expensive, tag-dependent hardware, while…
cs.CV2025
Generating on Generated: An Approach Towards Self-Evolving Diffusion Models
Xulu Zhang, Xiaoyong Wei, Jinlin Wu +4
Recursive Self-Improvement (RSI) enables intelligence systems to autonomously refine their capabilities. This paper explores the application of RSI in text-to-image diffusion model…
cs.CV2025
PolySmart @ TRECVid 2024 Video Captioning (VTT)
Jiaxin Wu, Wengyu Zhang, Xiao-Yong Wei +1
In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA…