From the 1 of 4 linked papers with an AI index.
4 papers
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Kaiqi Liu, Yunyao Mao, Ziqi Cai +8
MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…
One-step Latent-free Image Generation with Pixel Mean Flows
Yiyang Lu, Susie Lu, Qiao Sun +6
Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent adva…
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
Shaocong Xu, Songlin Wei, Qizhe Wei +12
Transparent objects remain notoriously hard for perception systems: refraction, reflection and transmission break the assumptions behind stereo, ToF and purely discriminative monoc…
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
Zheng Geng, Nan Wang, Shaocong Xu +5
Estimating the 6D pose of arbitrary unseen objects from a single reference image is critical for robotics operating in the long-tail of real-world instances. However, this setting…