3 papers
cs.CV2025
Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
Lakshita Agarwal, Bindu Verma
Understanding and analyzing video actions are essential for producing insightful and contextualized descriptions, especially for video-based applications like intelligent monitorin…
eess.IV2025
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism
Lakshita Agarwal, Bindu Verma
The examination of chest X-ray images is a crucial component in detecting various thoracic illnesses. This study introduces a new image description generation model that integrates…
cs.CV2025
Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism
Lakshita Agarwal, Bindu Verma
Image description generation is essential for accessibility and AI understanding of visual content. Recent advancements in deep learning have significantly improved natural languag…