#diffusion models
115 resultsTPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models
Taewon Kang, Matthias Zwicker
The paper introduces Temporal Prior Decoupling (TPD), a training‑free method that restores suppressed late‑segment information during diffusion sampling for text‑to‑video models, i…
Explicit Layer Modeling for Video Object Insertion and Layer Decomposition
Kyujin Han, Seungjoo Shin, Sunghyun Cho
The paper presents TriLayer, a large triplet video dataset with explicit foreground, background, and composite layers, and introduces DBL-Diffusion, a dual‑branch diffusion model t…
WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models
Jiamin Xu, Cong Wang, Zheng Dong +4
The paper introduces WildShadowRemover, a system that fine‑tunes a pretrained video diffusion model to remove shadows from real‑world videos while preserving fine details and tempo…
LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models
Bowen Chen, Shreshth Saini, Balu Adsumilli +1
LumaGuide is a training‑free framework that steers the sampling of pretrained diffusion models by matching target luminance distributions, enabling high dynamic range (HDR) image g…
Two2Four: Generative Quadruped Puppeteering from Human Motion
Fatemeh Zargarbashi, Zehong Qiu, Dhruv Agrawal +4
The paper introduces a two-stage generative diffusion framework that automatically converts ordinary human motion into realistic, controllable quadruped animations for virtual prod…
$Ï\mathbf{R}^2$: Reactive Real-time Flow Policies
Sungjae Park, Shubham Tulsiani
The paper introduces πR², a method that makes large pretrained manipulation policies reactive and real-time by separating fast proprioceptive inputs from slower vision-language inp…
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
Krzysztof Adamkiewicz, Brian Bernhard Moser, Stanislav Frolov +3
The paper evaluates modern text-to-image diffusion models as sources of synthetic training data and finds that, despite higher visual quality, newer models produce less diverse ima…
Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation
Qicheng Zhao, Qi Sun, Zheyu Yan
The paper presents Seer, a training‑free approach that detects the true end of generated sequences in diffusion multimodal large language models by monitoring MLP activation sparsi…
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Xinyu Liu, Shihao Li, Weihong Lin +10
The paper introduces ReBind, a framework that uses structured instructions with explicit reference tokens to improve multi‑reference image‑conditioned video editing, enabling preci…
Advanced Image Generation: Negative Prompt Optimization and Latent Classifier Guidance
Vaddi Charan Sai Nandan Reddy, Harini B, Chandana M S
The paper introduces a system that automatically creates optimized negative prompts using a fine‑tuned LLM and guides Stable Diffusion with a latent‑space CNN‑RNN classifier to red…
Avoiding Dilution: Using Diffusion and Vision Transformers to resolve Majorana Features in Nanowires at High Temperature
Jacob R. Taylor, Haining Pan, Jay D. Sau +1
The paper shows that neural networks based on diffusion-inspired U‑Net transformers and vision transformers can reconstruct low‑temperature conductance and predict topological visi…
Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation
Emanuele Colonna, Moises Diaz, Gennaro Vessio +2
The paper presents PIDiffSign, a diffusion-based model that generates 3D sign language motion from spoken language while enforcing anatomical constraints to ensure realistic and bi…
Rare Concept Generation via Counterfactual Inference in Diffusion Models
Zhengyuan Jiang, Haipeng Liu, Meng Wang +1
The paper introduces CI-Diff, a diffusion model that uses counterfactual causal inference to reduce common‑knowledge bias and better generate images of rare concepts described by t…
TAMF-VTON: Texture-Aware Mask-Free Virtual Try-On via High-Fidelity Image Synthesis
Jie Wang, Qian He, Gaofeng He +2
The paper introduces TAMF-VTON, a diffusion-based virtual try‑on system that works without segmentation masks and can apply multiple garments while preserving fine texture details,…
CoDi -- an exemplar-conditioned diffusion model for low-shot counting
Grega Å uÅ¡tar, Jer Pelhan, Alan LukežiÄ +1
CoDi is a latent diffusion-based model that uses exemplar-conditioned conditioning to generate high-quality density maps for low-shot object counting, enabling accurate object loca…
VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents
Marcus Hoerger, Rishikesh Joshi, Rahul Shome +2
VOiLA learns task‑agnostic POMDP transition and observation models with conditional diffusion networks, distills them into fast feed‑forward generators, and integrates them with a…
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
Yushi Huang, Xiangxin Zhou, Jun Zhang +2
The paper introduces MeanFlowNFT, a method that applies reinforcement‑learning based reward optimization to MeanFlow generators by learning an instantaneous‑velocity predictor whil…
MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
Yinhan Zhang, Dingwei Tan, Dinwei Tan +4
The paper introduces MagicPrompt, a lightweight method that uses attention-embedded soft prompts and dual-space reward feedback to fine‑tune large video diffusion models with less…
WanSong v1.0 Technical Report
Binghui Chen, Pandeng Li, Yu Liu +1
The paper introduces WanSong, a diffusion‑based model that can directly generate high‑fidelity, multilingual songs up to five minutes long, outputting separate vocal and background…
VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios
Kailin Lyu, Long Xiao, Jianing Zeng +3
The paper presents VQ-Touch, a framework that efficiently generates high‑fidelity tactile images across different sensors and scenarios using a VQ‑GAN based representation and a di…
Introspective Attention Modulation for Safe Text-to-Image Generation
Basim Azam, Hossein Rahmani, Naveed Akhtar
The paper proposes a method that monitors and adjusts the attention mechanisms of text‑to‑image diffusion models at inference time to prevent the generation of unsafe content while…
Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion
Wenqi Si, Gongyang Li, Shixiang Shi +1
The paper proposes a weakly‑supervised RGB‑D salient object detection framework that expands sparse scribble labels into dense pseudo annotations using the Segment Anything Model a…
RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
Tianyuan Qu, Lei Ke, Xiaohang Zhan +6
The paper presents RePlan, a framework that first reasons about natural‑language instructions to identify specific image regions and then edits those regions using a diffusion mode…
UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation
Junjie Lu, Xinyao Qin, Yuhua Jiang +6
The paper introduces UniSteer, a framework that converts human corrective actions into noise targets to guide a lightweight noise-prediction actor while simultaneously training it…