From the 1 of 9 linked papers with an AI index.
9 papers
SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
Zheng Liu, Zijian He, Huiguo He +5
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span differ…
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Mingju Gao, Jingkai Zhou, Kun Gai +2
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-match…
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
Yufei Cai, Xuesong Niu, Hao Lu +3
MetaView is a diffusion-based framework that generates novel views from a single image by combining implicit geometry priors with metric depth cues, enabling large viewpoint change…
NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning
Tianlin Pan, Lianyu Pang, Cheng Da +4
Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward…
Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions
Bo Zhao, Kairui Guo, Runnan Du +6
Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many seemingly simple cases. We obs…
MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training
Lianyu Pang, Tianlin Pan, Cheng Da +5
Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning intermediate diffusion featu…