activity
20242026
collaborators

6 papers

cs.CV2026

Attention-Steered Vision-Language Models for Sign Language Translation

Meibo Hu, Guohao Sun, Annemarie D. Ross +2

Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we…

cs.CV2026

Information-Regularized Attention for Visual-Centric Reasoning

Guohao Sun, Xiaofang Wang, Yash Patel +3

Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…

cs.AI2025

Latent Chain-of-Thought for Visual Reasoning

Guohao Sun, Hang Hua, Jian Wang +5

Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…

cs.IR2024

RaCT: Ranking-aware Chain-of-Thought Optimization for LLMs

Haowei Liu, Xuyang Wu, Guohao Sun +2

In information retrieval, large language models (LLMs) have demonstrated remarkable potential in text reranking tasks by leveraging their sophisticated natural language understandi…

cs.CV2024

Prototypical Transformer as Unified Motion Learners

Cheng Han, Yawen Lu, Guohao Sun +9

In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoForme…

cs.CV2024

STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering

Guohao Sun, Can Qin, Huazhu Fu +2

Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medica…