1 paper
Meibo Hu, Guohao Sun, Annemarie D. Ross +2
Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we…