4 papers
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
Haixu Liu, Penghao Jiang, Zerui Tao +2
Predicting plant species composition in specific spatiotemporal contexts plays an important role in biodiversity management and conservation, as well as in improving species identi…
Google is all you need: Semi-Supervised Transfer Learning Strategy For Light Multimodal Multi-Task Classification Model
Haixu Liu, Penghao Jiang, Zerui Tao
As the volume of digital image data increases, the effectiveness of image classification intensifies. This study introduces a robust multi-label classification system designed to a…
Multi-Modal Video Feature Extraction for Popularity Prediction
Haixu Liu, Wenning Wang, Haoxiang Zheng +4
This work aims to predict the popularity of short videos using the videos themselves and their related features. Popularity is measured by four key engagement metrics: view count,…
Unveiling and Controlling Anomalous Attention Distribution in Transformers
Ruiqing Yan, Xingbo Du, Haoyu Deng +7
With the advent of large models based on the Transformer architecture, researchers have observed an anomalous phenomenon in the Attention mechanism--there is a very high attention…