collaborators

6 papers

cs.CV2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

Xinyu Zhang, Yuxuan Dong, Lingling Zhang +5

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrele…

cs.CV2025

Learning to Generate Long-term Future Narrations Describing Activities of Daily Living

Ramanathan Rajendiran, Debaditya Roy, Basura Fernando

Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich…

cs.AI2025

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

Xinyu Zhang, Yuxuan Dong, Yanrui Wu +6

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based re…

cs.LG2024

Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios

Shantanu Jaiswal, Debaditya Roy, Basura Fernando +1

Complex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the imm…

cs.CV2024

Diagram-Driven Course Questions Generation

Xinyu Zhang, Lingling Zhang, Yanrui Wu +6

Visual Question Generation (VQG) research focuses predominantly on natural images while neglecting the diagram, which is a critical component in educational materials. To meet the…

cs.LG2024

FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data

Binqian Xu, Xiangbo Shu, Haiyang Mei +3

Multimodal Large Language Models (MLLMs) have made significant advancements, demonstrating powerful capabilities in processing and understanding multimodal data. Fine-tuning MLLMs…