5 papers
EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
Umar Khalid, Kashif Munir, Hasan Iqbal +6
Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content bu…
PSF-4D: A Progressive Sampling Framework for View Consistent 4D Editing
Hasan Iqbal, Nazmul Karim, Umar Khalid +4
Instruction-guided generative models, especially those using text-to-image (T2I) and text-to-video (T2V) diffusion frameworks, have advanced the field of content editing in recent…
3DEgo: 3D Editing on the Go!
Umar Khalid, Hasan Iqbal, Azib Farooq +2
We introduce 3DEgo to address a novel problem of directly synthesizing photorealistic 3D scenes from monocular videos guided by textual prompts. Conventional methods construct a te…
Free-Editor: Zero-shot Text-driven 3D Scene Editing
Nazmul Karim, Hasan Iqbal, Umar Khalid +2
Text-to-Image (T2I) diffusion models have recently gained traction for their versatility and user-friendliness in 2D content generation and editing. However, training a diffusion m…
LatentEditor: Text Driven Local Editing of 3D Scenes
Umar Khalid, Hasan Iqbal, Nazmul Karim +2
While neural fields have made significant strides in view synthesis and scene reconstruction, editing them poses a formidable challenge due to their implicit encoding of geometry a…