6 papers
Attention-Steered Vision-Language Models for Sign Language Translation
Meibo Hu, Guohao Sun, Annemarie D. Ross +2
Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we…
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, Yash Patel +3
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…
Latent Chain-of-Thought for Visual Reasoning
Guohao Sun, Hang Hua, Jian Wang +5
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…
RaCT: Ranking-aware Chain-of-Thought Optimization for LLMs
Haowei Liu, Xuyang Wu, Guohao Sun +2
In information retrieval, large language models (LLMs) have demonstrated remarkable potential in text reranking tasks by leveraging their sophisticated natural language understandi…
Prototypical Transformer as Unified Motion Learners
Cheng Han, Yawen Lu, Guohao Sun +9
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoForme…
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
Guohao Sun, Can Qin, Huazhu Fu +2
Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medica…