NewEvery arXiv paper, its researchers & institutions — mapped.
the archive

#diffusion models

115 results
cs.CV2026

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models

Taewon Kang, Matthias Zwicker

The paper introduces Temporal Prior Decoupling (TPD), a training‑free method that restores suppressed late‑segment information during diffusion sampling for text‑to‑video models, i…

#text-to-video generation#diffusion models#temporal coherence#classifier-free guidance
cs.CV2026

Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

Kyujin Han, Seungjoo Shin, Sunghyun Cho

The paper presents TriLayer, a large triplet video dataset with explicit foreground, background, and composite layers, and introduces DBL-Diffusion, a dual‑branch diffusion model t…

#video editing#layered video representation#object insertion#video decomposition
cs.CV2026

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

Jiamin Xu, Cong Wang, Zheng Dong +4

The paper introduces WildShadowRemover, a system that fine‑tunes a pretrained video diffusion model to remove shadows from real‑world videos while preserving fine details and tempo…

#video shadow removal#diffusion models#detail preservation#depth priors
cs.CV2026

LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models

Bowen Chen, Shreshth Saini, Balu Adsumilli +1

LumaGuide is a training‑free framework that steers the sampling of pretrained diffusion models by matching target luminance distributions, enabling high dynamic range (HDR) image g…

#diffusion models#hdr imaging#distribution shaping#energy-based guidance
cs.GR2026

Two2Four: Generative Quadruped Puppeteering from Human Motion

Fatemeh Zargarbashi, Zehong Qiu, Dhruv Agrawal +4

The paper introduces a two-stage generative diffusion framework that automatically converts ordinary human motion into realistic, controllable quadruped animations for virtual prod…

#quadruped animation#motion retargeting#diffusion models#human motion capture
cs.RO2026

$π\mathbf{R}^2$: Reactive Real-time Flow Policies

Sungjae Park, Shubham Tulsiani

The paper introduces πR², a method that makes large pretrained manipulation policies reactive and real-time by separating fast proprioceptive inputs from slower vision-language inp…

#manipulation#real-time control#diffusion models#reactive policies
cs.CV2026

When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators

Krzysztof Adamkiewicz, Brian Bernhard Moser, Stanislav Frolov +3

The paper evaluates modern text-to-image diffusion models as sources of synthetic training data and finds that, despite higher visual quality, newer models produce less diverse ima…

#text-to-image generation#synthetic data#image classification#data diversity
cs.AI2026

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation

Qicheng Zhao, Qi Sun, Zheyu Yan

The paper presents Seer, a training‑free approach that detects the true end of generated sequences in diffusion multimodal large language models by monitoring MLP activation sparsi…

#diffusion models#multimodal large language models#inference acceleration#sequence truncation
cs.CV2026

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Xinyu Liu, Shihao Li, Weihong Lin +10

The paper introduces ReBind, a framework that uses structured instructions with explicit reference tokens to improve multi‑reference image‑conditioned video editing, enabling preci…

#multi-reference video editing#diffusion models#structured instructions#reference tokens
cs.CV2026

Advanced Image Generation: Negative Prompt Optimization and Latent Classifier Guidance

Vaddi Charan Sai Nandan Reddy, Harini B, Chandana M S

The paper introduces a system that automatically creates optimized negative prompts using a fine‑tuned LLM and guides Stable Diffusion with a latent‑space CNN‑RNN classifier to red…

#image generation#diffusion models#negative prompting#latent guidance
cond-mat.mes-hall2026

Avoiding Dilution: Using Diffusion and Vision Transformers to resolve Majorana Features in Nanowires at High Temperature

Jacob R. Taylor, Haining Pan, Jay D. Sau +1

The paper shows that neural networks based on diffusion-inspired U‑Net transformers and vision transformers can reconstruct low‑temperature conductance and predict topological visi…

#majorana zero modes#nanowire devices#high-temperature screening#machine learning
cs.CV2026

Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation

Emanuele Colonna, Moises Diaz, Gennaro Vessio +2

The paper presents PIDiffSign, a diffusion-based model that generates 3D sign language motion from spoken language while enforcing anatomical constraints to ensure realistic and bi…

#sign language generation#diffusion models#physics-informed learning#3d pose estimation
cs.CV2026

Rare Concept Generation via Counterfactual Inference in Diffusion Models

Zhengyuan Jiang, Haipeng Liu, Meng Wang +1

The paper introduces CI-Diff, a diffusion model that uses counterfactual causal inference to reduce common‑knowledge bias and better generate images of rare concepts described by t…

#rare concept generation#diffusion models#counterfactual inference#causal inference
cs.CV2026

TAMF-VTON: Texture-Aware Mask-Free Virtual Try-On via High-Fidelity Image Synthesis

Jie Wang, Qian He, Gaofeng He +2

The paper introduces TAMF-VTON, a diffusion-based virtual try‑on system that works without segmentation masks and can apply multiple garments while preserving fine texture details,…

#virtual try-on#texture preservation#mask-free synthesis#multi-garment composition
cs.CV2026

CoDi -- an exemplar-conditioned diffusion model for low-shot counting

Grega Šuštar, Jer Pelhan, Alan Lukežič +1

CoDi is a latent diffusion-based model that uses exemplar-conditioned conditioning to generate high-quality density maps for low-shot object counting, enabling accurate object loca…

#low-shot counting#diffusion models#density estimation#exemplar conditioning
cs.RO2026

VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents

Marcus Hoerger, Rishikesh Joshi, Rahul Shome +2

VOiLA learns task‑agnostic POMDP transition and observation models with conditional diffusion networks, distills them into fast feed‑forward generators, and integrates them with a…

#pomdp planning#diffusion models#online planning#belief updates
cs.CV2026

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

Yushi Huang, Xiangxin Zhou, Jun Zhang +2

The paper introduces MeanFlowNFT, a method that applies reinforcement‑learning based reward optimization to MeanFlow generators by learning an instantaneous‑velocity predictor whil…

#meanflow generators#diffusion models#reinforcement learning#few-step sampling
cs.CV2026

MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

Yinhan Zhang, Dingwei Tan, Dinwei Tan +4

The paper introduces MagicPrompt, a lightweight method that uses attention-embedded soft prompts and dual-space reward feedback to fine‑tune large video diffusion models with less…

#video generation#diffusion models#prompt tuning#parameter-efficient fine-tuning
eess.AS2026

WanSong v1.0 Technical Report

Binghui Chen, Pandeng Li, Yu Liu +1

The paper introduces WanSong, a diffusion‑based model that can directly generate high‑fidelity, multilingual songs up to five minutes long, outputting separate vocal and background…

#music generation#diffusion models#long-form audio#multilingual songs
cs.CV2026

VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios

Kailin Lyu, Long Xiao, Jianing Zeng +3

The paper presents VQ-Touch, a framework that efficiently generates high‑fidelity tactile images across different sensors and scenarios using a VQ‑GAN based representation and a di…

#tactile sensing#image generation#cross‑sensor generalization#few‑shot learning
cs.CV2026

Introspective Attention Modulation for Safe Text-to-Image Generation

Basim Azam, Hossein Rahmani, Naveed Akhtar

The paper proposes a method that monitors and adjusts the attention mechanisms of text‑to‑image diffusion models at inference time to prevent the generation of unsafe content while…

#safe text-to-image generation#attention modulation#diffusion models#inference-time control
cs.CV2026

Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion

Wenqi Si, Gongyang Li, Shixiang Shi +1

The paper proposes a weakly‑supervised RGB‑D salient object detection framework that expands sparse scribble labels into dense pseudo annotations using the Segment Anything Model a…

#salient object detection#rgb-d data#weak supervision#diffusion models
cs.CV2026

RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing

Tianyuan Qu, Lei Ke, Xiaohang Zhan +6

The paper presents RePlan, a framework that first reasons about natural‑language instructions to identify specific image regions and then edits those regions using a diffusion mode…

#instruction-based image editing#region planning#vision-language reasoning#diffusion models
cs.RO2026

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation

Junjie Lu, Xinyao Qin, Yuhua Jiang +6

The paper introduces UniSteer, a framework that converts human corrective actions into noise targets to guide a lightweight noise-prediction actor while simultaneously training it…

#vision-language-action#diffusion models#reinforcement learning#human-in-the-loop