activity
20172025
most citedStoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation

5 citations · 9 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CL2025

REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation

Dong-Ho Lee, Adyasha Maharana, Jay Pujara +2

Long-term, open-domain dialogue capabilities are essential for chatbots aiming to recall past interactions and demonstrate emotional intelligence (EI). Yet, most existing research…

cs.LG2024

Adapt-: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection

Adyasha Maharana, Jaehong Yoon, Tianlong Chen +1

Visual instruction datasets from various distributors are released at different times and often contain a significant number of semantically redundant text-image pairs, depending o…

cs.LG2023

Debiasing Multimodal Models via Causal Information Minimization

Vaidehi Patil, Adyasha Maharana, Mohit Bansal

Most existing debiasing methods for multimodal models, including causal intervention and inference methods, utilize approximate heuristics to represent the biases, such as shallow…

cs.CV20225 cited

StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation

Adyasha Maharana, Darryl Hannan, Mohit Bansal

Recent advances in text-to-image synthesis have led to large pretrained transformers with excellent capabilities to generate visualizations from a given text. However, these models…

cs.CL2021

Integrating Visuospatial, Linguistic and Commonsense Structure into Story Visualization

Adyasha Maharana, Mohit Bansal

While much research has been done in text-to-image synthesis, little work has been done to explore the usage of linguistic structure of the input text. Such information is even mor…

cs.CL20211 cited

Improving Generation and Evaluation of Visual Stories via Semantic Consistency

Adyasha Maharana, Darryl Hannan, Mohit Bansal

Story visualization is an under-explored task that falls at the intersection of many important research directions in both computer vision and natural language processing. In this…