2 papers
cs.CV2025
EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
Umar Khalid, Kashif Munir, Hasan Iqbal +6
Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content bu…
cs.CV2025
PSF-4D: A Progressive Sampling Framework for View Consistent 4D Editing
Hasan Iqbal, Nazmul Karim, Umar Khalid +4
Instruction-guided generative models, especially those using text-to-image (T2I) and text-to-video (T2V) diffusion frameworks, have advanced the field of content editing in recent…