1 citations · 1 across the 3 of their papers we have counts for
5 papers
MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills
Niladri Shekhar Dutt, Duygu Ceylan, Niloy J. Mitra
Retouching is an essential task in post-manipulation of raw photographs. Generative editing, guided by text or strokes, provides a new tool accessible to users but can easily chang…
JOG3R: Towards 3D-Consistent Video Generators
Chun-Hao Paul Huang, Niloy Mitra, Hyeonho Jeong +2
Emergent capabilities of image generators have led to many impactful zero- or few-shot applications. Inspired by this success, we investigate whether video generators similarly exh…
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
Souhaib Attaiki, Paul Guerrero, Duygu Ceylan +2
We train a feed-forward text-to-3D diffusion generator for human characters using only single-view 2D data for supervision. Existing 3D generative models cannot yet match the fidel…
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye +2
While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across…
Motion Modes: What Could Happen Next?
Karran Pandey, Matheus Gadelha, Yannick Hold-Geoffroy +3
Predicting diverse object motions from a single static image remains challenging, as current video generation models often entangle object movement with camera motion and other sce…