Showing 2024Show all
2 papers · 1 filter
cs.CV2024
EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
Umar Khalid, Kashif Munir, Hasan Iqbal +6
Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content bu…
cs.CV2024
3DEgo: 3D Editing on the Go!
Umar Khalid, Hasan Iqbal, Azib Farooq +2
We introduce 3DEgo to address a novel problem of directly synthesizing photorealistic 3D scenes from monocular videos guided by textual prompts. Conventional methods construct a te…