2 papers
cs.CV2025
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
Liu Jing, Amirul Rahman
Large Vision-Language Models (LVLMs) have shown remarkable progress in various multimodal tasks, yet they often struggle with complex visual reasoning that requires multi-step infe…
cs.CV2024
Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction
Liu Jing, Amirul Rahman
Semantic location prediction from multimodal social media posts is a critical task with applications in personalized services and human mobility analysis. This paper introduces \te…