24 citations · 46 across the 17 of their papers we have counts for
7 papers · 1 filter
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
Daizong Liu, Mingyu Yang, Xiaoye Qu +3
With the significant development of large models in recent years, Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal u…
Density-Insensitive Unsupervised Domain Adaption on 3D Object Detection
Qianjiang Hu, Daizong Liu, Wei Hu
3D object detection from point clouds is crucial in safety-critical autonomous driving. Although many works have made great efforts and achieved significant progress on this task,…
Neural Capture of Animatable 3D Human from Monocular Video
Gusi Te, Xiu Li, Xiao Li +3
We present a novel paradigm of building an animatable 3D human representation from a monocular video input, such that it can be rendered in any unseen poses and views. Our method i…
Reducing the Vision and Language Bias for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Wei Hu
Temporal sentence grounding (TSG) is an important yet challenging task in multimedia information retrieval. Although previous TSG methods have achieved decent performance, they ten…
Skimming, Locating, then Perusing: A Human-Like Framework for Natural Language Video Localization
Daizong Liu, Wei Hu
This paper addresses the problem of natural language video localization (NLVL). Almost all existing works follow the "only look once" framework that exploits a single model to dire…
Unsupervised Manga Character Re-identification via Face-body and Spatial-temporal Associated Clustering
Zhimin Zhang, Zheng Wang, Wei Hu
In the past few years, there has been a dramatic growth in e-manga (electronic Japanese-style comics). Faced with the booming demand for manga research and the large amount of unla…