2 papers
cs.LG2024
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
Jinyin Wang, Haijing Zhang, Yihao Zhong +3
Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia infor…
cs.CV2024
Research on Image Recognition Technology Based on Multimodal Deep Learning
Jinyin Wang, Xingchen Li, Yixuan Jin +3
This project investigates the human multi-modal behavior identification algorithm utilizing deep neural networks. According to the characteristics of different modal information, d…