29 citations · 33 across the 3 of their papers we have counts for
5 papers · 1 filter
Multimodal Prompt Alignment for Facial Expression Recognition
Fuyan Ma, Yiran He, Bin Sun +1
Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their success, current VLM-based facial e…
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
Ziyu Ma, Shutao Li, Bin Sun +3
Knowledge-based visual question answering (VQA) requires world knowledge beyond the image for accurate answer. Recently, instead of extra knowledge bases, a large language model (L…
LOGO-Former: Local-Global Spatio-Temporal Transformer for Dynamic Facial Expression Recognition
Fuyan Ma, Bin Sun, Shutao Li
Previous methods for dynamic facial expression recognition (DFER) in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range…
Spatio-Temporal Transformer for Dynamic Facial Expression Recognition in the Wild
Fuyan Ma, Bin Sun, Shutao Li
Previous methods for dynamic facial expression in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range dependencies in vi…
Hybrid Mutimodal Fusion for Dimensional Emotion Recognition
Ziyu Ma, Fuyan Ma, Bin Sun +1
In this paper, we extensively present our solutions for the MuSe-Stress sub-challenge and the MuSe-Physio sub-challenge of Multimodal Sentiment Challenge (MuSe) 2021. The goal of M…