2 papers
cs.CV2026
Dual-Stream Collaborative Transformer for Image Captioning
Jun Wan, Jun Liu, Zhihui lai +1
Current region feature-based image captioning methods have progressed rapidly and achieved remarkable performance. However, they are still prone to generating irrelevant descriptio…
cs.CV2024
Precise Facial Landmark Detection by Dynamic Semantic Aggregation Transformer
Jun Wan, He Liu, Yujia Wu +3
At present, deep neural network methods have played a dominant role in face alignment field. However, they generally use predefined network structures to predict landmarks, which t…