activity
20222024
most citedCollaborative Three-Stream Transformers for Video Captioning

8 citations · 30 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

17 papers · 1 filter

cs.CV20251 cited

High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion

Libo Zhang, Yongsheng Yu, Jiali Yao +1

Generative Adversarial Network (GAN) inversion have demonstrated excellent performance in image inpainting that aims to restore lost or damaged image texture using its unmasked con…

cs.CV2024

PRTGaussian: Efficient Relighting Using 3D Gaussians with Precomputed Radiance Transfer

Libo Zhang, Yuxuan Han, Wenbin Lin +2

We present PRTGaussian, a realtime relightable novel-view synthesis method made possible by combining 3D Gaussians and Precomputed Radiance Transfer (PRT). By fitting relightable G…

cs.CV20241 cited

LaMOT: Language-Guided Multi-Object Tracking

Yunhao Li, Xiaoqiong Liu, Luke Liu +2

Vision-Language MOT is a crucial tracking problem and has drawn increasing attention recently. It aims to track objects based on human language commands, replacing the traditional…

cs.CV20244 cited

VastTrack: Vast Category Visual Object Tracking

Liang Peng, Junyuan Gao, Xinran Liu +5

In this paper, we introduce a novel benchmark, dubbed VastTrack, towards facilitating the development of more general visual tracking via encompassing abundant classes and videos.…

cs.CV20244 cited

CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition

Jinzhi Zheng, Ruyi Ji, Libo Zhang +2

Scene text recognition, as a cross-modal task involving vision and text, is an important research topic in computer vision. Most existing methods use language models to extract sem…

cs.CV2024

Text Region Multiple Information Perception Network for Scene Text Detection

Jinzhi Zheng, Libo Zhang, Yanjun Wu +1

Segmentation-based scene text detection algorithms can handle arbitrary shape scene texts and have strong robustness and adaptability, so it has attracted wide attention. Existing…