8 papers · 1 filter
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
Dongxu Ge, Shansong Liu, Cheng Gong +3
As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted incre…
Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection
Yuchu Jiang, Jiaming Chu, Jian Zhao +5
The proliferation of generative models has raised serious concerns about visual content forgery. Existing deepfake detection methods primarily target either image-level classificat…
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +5
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion mo…
NFIG: Multi-Scale Autoregressive Image Generation via Frequency Ordering
Zhihao Huang, Xi Qiu, Yukuo Ma +5
Autoregressive models have achieved significant success in image generation. However, unlike the inherent hierarchical structure of image information in the spectral domain, standa…
Conditional Video Generation for High-Efficiency Video Compression
Fangqiu Yi, Jingyu Xu, Jiawei Shao +2
Perceptual studies demonstrate that conditional diffusion models excel at reconstructing video content aligned with human visual perception. Building on this insight, we propose a…
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
Chenyou Fan, Fangzheng Yan, Chenjia Bai +4
Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existin…