3 papers
cs.CV2025
Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams
Minghao Zhou, Rafael Souza, Yaqian Hu +1
Large Language Models (LLMs) and their multimodal variants (LVLMs) hold immense promise for scientific and engineering applications, particularly in processing visual information l…
cs.CV2025
Knowledge-Augmented Language Models Interpreting Structured Chest X-Ray Findings
Alexander Davis, Rafael Souza, Jia-Hao Lim
Automated interpretation of chest X-rays (CXR) is a critical task with the potential to significantly improve clinical workflow and patient care. While recent advances in multimoda…
cs.CV2024
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
Rafael Souza, Jia-Hao Lim, Alexander Davis
Temporal reasoning is a critical challenge in video-language understanding, as it requires models to align semantic concepts consistently across time. While existing large vision-l…