3 citations · 5 across the 18 of their papers we have counts for
8 papers · 1 filter
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
Dongxu Ge, Shansong Liu, Cheng Gong +3
As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted incre…
Conditional Video Generation for High-Efficiency Video Compression
Fangqiu Yi, Jingyu Xu, Jiawei Shao +2
Perceptual studies demonstrate that conditional diffusion models excel at reconstructing video content aligned with human visual perception. Building on this insight, we propose a…
Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection
Yuchu Jiang, Jiaming Chu, Jian Zhao +5
The proliferation of generative models has raised serious concerns about visual content forgery. Existing deepfake detection methods primarily target either image-level classificat…
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
Chenyou Fan, Fangzheng Yan, Chenjia Bai +4
Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existin…
Metric-Solver: Sliding Anchored Metric Depth Estimation from a Single Image
Tao Wen, Jiepeng Wang, Yabo Chen +3
Accurate and generalizable metric depth estimation is crucial for various computer vision applications but remains challenging due to the diverse depth scales encountered in indoor…
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +5
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion mo…