From the 1 of 4 linked papers with an AI index.
4 papers
Fine-grained CLIP fine-tuning with self-annotated region alignment
Chenyang Zhao, Wei Lin, Antoni B. Chan +1
The paper proposes SFF-CLIP, a fine-tuning approach that uses only image-text pairs to align region features with phrase concepts via text-specific heat maps, improving CLIP's fine…
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
Chenyang Zhao, Kun Wang, Janet H. Hsiao +1
Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is…
Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting
Wei Lin, Chenyang Zhao, Antoni B. Chan
Point detection has been developed to locate pedestrians in crowded scenes by training a counter through a point-to-point (P2P) supervision scheme. Despite its excellent localizati…
Density-based Object Detection in Crowded Scenes
Chenyang Zhao, Jia Wan, Antoni B. Chan
Compared with the generic scenes, crowded scenes contain highly-overlapped instances, which result in: 1) more ambiguous anchors during training of object detectors, and 2) more pr…