3 papers
cs.CV2026
ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features
Jay Mahajan, Chang Liu, Rauf Makharov +3
We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Transformer (DiT) models via their internal feature representations. Rea…
cs.AI2023
MineObserver 2.0: A Deep Learning & In-Game Framework for Assessing Natural Language Descriptions of Minecraft Imagery
Jay Mahajan, Samuel Hum, Jack Henhapl +5
MineObserver 2.0 is an AI framework that uses Computer Vision and Natural Language Processing for assessing the accuracy of learner-generated descriptions of Minecraft images that…
cs.CV2023
Street TryOn: Learning In-the-Wild Virtual Try-On from Unpaired Person Images
Aiyu Cui, Jay Mahajan, Viraj Shah +3
Most virtual try-on research is motivated to serve the fashion business by generating images to demonstrate garments on studio models at a lower cost. However, virtual try-on shoul…