V2VCrafter: Consistent Street-View Image Generation Across Vehicles
arXiv:2605.29471
Abstract
Connected and autonomous driving (CAD) systems leverage vehicle-to-vehicle (V2V) communication for multi-agent collaborative perception, yet remain constrained by scarce annotated real-world V2V datasets and limited generalization across diverse driving conditions. While image generation offers a feasible solution for data augmentation, existing single-vehicle multi-view generation frameworks face two key challenges in multi-agent settings: (1) the expanded learning objective degrades generation quality, and (2) dynamic inter-agent variation hinders consistency modeling for physical attributes (e.g., color, category) of jointly observed objects. To bridge this gap, we propose V2VCrafter, the first framework for generating controllable and realistic multi-view driving images across vehicles. For effective learning, we develop a progressive multi-agent diffusion model based on a single-agent backbone, using neighboring agents' latent states to progressively guide single-to-multi-agent generation. To address cross-vehicle inconsistency, we further propose a cross-agent attention module that leverages a collaboration view graph and learnable jointly observed object representations to model dynamic cross-vehicle camera view relationships. Experiments on real-world V2X-Real dataset show that V2VCrafter generates high-fidelity, controllable, and consistent street views across vehicles, thereby effectively enhancing downstream collaborative 3D object detection tasks.