activity
20222024
most citedFrom Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion Retrieval

12 citations · 22 across the 10 of their papers we have counts for

collaborators

9 papers

cs.CV2023

Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action Understanding

Shengkai Sun, Daizong Liu, Jianfeng Dong +5

Unsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models, then integrate…

cs.CV2023

Dense Object Grounding in 3D Scenes

Wencan Huang, Daizong Liu, Wei Hu

Localizing objects in 3D scenes according to the semantics of a given natural language is a fundamental yet important task in the field of multimedia understanding, which benefits…

cs.CV20232 cited

3DHacker: Spectrum-based Decision Boundary Generation for Hard-label 3D Point Cloud Attack

Yunbo Tao, Daizong Liu, Pan Zhou +3

With the maturity of depth sensors, the vulnerability of 3D point cloud models has received increasing attention in various applications such as autonomous driving and robot naviga…

cs.CV202312 cited

From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion Retrieval

Jianfeng Dong, Xiaoman Peng, Zhe Ma +5

Attribute-specific fashion retrieval (ASFR) is a challenging information retrieval task, which has attracted increasing attention in recent years. Different from traditional fashio…

cs.CV2023

Jointly Visual- and Semantic-Aware Graph Memory Networks for Temporal Sentence Localization in Videos

Daizong Liu, Pan Zhou

Temporal sentence localization in videos (TSLV) aims to retrieve the most interested segment in an untrimmed video according to a given sentence query. However, almost of existing…

cs.CV2023

Tracking Objects and Activities with Attention for Temporal Sentence Grounding

Zeyu Xiong, Daizong Liu, Pan Zhou +1

Temporal sentence grounding (TSG) aims to localize the temporal segment which is semantically aligned with a natural language query in an untrimmed video.Most existing methods extr…