4 papers · 1 filter
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
Rohan Wadhawan, Hritik Bansal, Kai-Wei Chang +1
Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual…
Multi-Attributed and Structured Text-to-Face Synthesis
Rohan Wadhawan, Tanuj Drall, Shubham Singh +1
Generative Adversarial Networks (GANs) have revolutionized image synthesis through many applications like face generation, photograph editing, and image super-resolution. Image syn…
Landmark-Aware and Part-based Ensemble Transfer Learning Network for Facial Expression Recognition from Static images
Rohan Wadhawan, Tapan K. Gandhi
Facial Expression Recognition from static images is a challenging problem in computer vision applications. Convolutional Neural Network (CNN), the state-of-the-art method for vario…
Intelligent Monitoring of Stress Induced by Water Deficiency in Plants using Deep Learning
Shiva Azimi, Rohan Wadhawan, Tapan K. Gandhi
In the recent decade, high-throughput plant phenotyping techniques, which combine non-invasive image analysis and machine learning, have been successfully applied to identify and q…