From the 1 of 10 linked papers with an AI index.
9 papers
Emotion Recognition in Sign Language Conversation
Yusong Wang, Keyu Mao, Takao Obi +2
The paper introduces emotion recognition in sign language conversations, presenting a new dataset (eJSL Dialog) of 1,920 video samples across 480 dialogues and evaluating various v…
BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming
Jiaxiang Liu, Tianxiang Hu, Juwei Guan +5
Recent advances in vision-language models (VLMs) such as CLIP have demonstrated strong generalization across natural-image domains. However, adapting these models to biomedical ima…
Toward Signing Activity Projection in Sign Language Interaction
Takao Obi, Wang Yusong, Koji Inoue +1
Social robots must interact robustly not only with users assumed by speech-centered systems but also with diverse users whose communication relies on different modalities, e.g., si…
ReverseEOL: Improving Training-free Text Embeddings via Text Reversal in Decoder-only LLMs
Ailiang Lin, Zhuoyun Li, Yusong Wang +3
Recent advances in Large Language Models (LLMs) have opened new avenues for generating training-free text embeddings. However, the causal attention in decoder-only LLMs prevents ea…
Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models
Xiao Liu, Jiaxiang Liu, Boci Peng +6
Vision Language Models adapt well to downstream tasks but are highly vulnerable to adversarial perturbations that disrupt cross-modal semantic alignment. Existing defenses are larg…
Exposing and Mitigating Temporal Attack in Deepfake Video Detection
Zheyuan Gu, Minghao Shao, Zhen Wang +4
While spatiotemporal deepfake detectors achieve high AUC, our experiments reveal their susceptibility to evasion attacks. These models tend to overfit on fragile temporal spectrum…