5 papers
Attention-Steered Vision-Language Models for Sign Language Translation
Meibo Hu, Guohao Sun, Annemarie D. Ross +2
Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we…
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, Yash Patel +3
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…
Latent Chain-of-Thought for Visual Reasoning
Guohao Sun, Hang Hua, Jian Wang +5
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…
RaCT: Ranking-aware Chain-of-Thought Optimization for LLMs
Haowei Liu, Xuyang Wu, Guohao Sun +2
In information retrieval, large language models (LLMs) have demonstrated remarkable potential in text reranking tasks by leveraging their sophisticated natural language understandi…
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
Guohao Sun, Can Qin, Huazhu Fu +2
Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medica…