most citedMonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.GR20251 cited

MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills

Niladri Shekhar Dutt, Duygu Ceylan, Niloy J. Mitra

Retouching is an essential task in post-manipulation of raw photographs. Generative editing, guided by text or strokes, provides a new tool accessible to users but can easily chang…

cs.CV2025

JOG3R: Towards 3D-Consistent Video Generators

Chun-Hao Paul Huang, Niloy Mitra, Hyeonho Jeong +2

Emergent capabilities of image generators have led to many impactful zero- or few-shot applications. Inspired by this success, we investigate whether video generators similarly exh…

cs.CV2024

GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space

Souhaib Attaiki, Paul Guerrero, Duygu Ceylan +2

We train a feed-forward text-to-3D diffusion generator for human characters using only single-view 2D data for supervision. Existing 3D generative models cannot yet match the fidel…

cs.CV2024

Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye +2

While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across…

cs.CV2024

Motion Modes: What Could Happen Next?

Karran Pandey, Matheus Gadelha, Yannick Hold-Geoffroy +3

Predicting diverse object motions from a single static image remains challenging, as current video generation models often entangle object movement with camera motion and other sce…