collaborators

10 papers

cs.CV2026

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

Zhennan Chen, Tianxing Shi, Pengcheng Xu +5

VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple obj…

cs.CV2026

Spiking Pyramid Wavelet Transformation for High-efficient and Low-energy Image Restoration

Chen Zhao, Xiantao Hu, Song Wu +5

Spiking neural networks (SNNs) have garnered significant interest in computer vision due to their potential for efficiency and biological inspiration. While spiking CNN-based metho…

cs.CV2026

TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On

Dingbao Shao, Song Wu, Shenyi Wang +9

Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limited. In this paper, we first i…

cs.RO2026

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement

Tianyidan Xie, Peiyu Wang, Yuyi Qian +7

Physics-aware symbolic simulation of 3D scenes is critical for robotics, embodied AI, and scientific computing, requiring models to understand natural language descriptions of phys…

cs.CV2026

A training-free framework for high-fidelity appearance transfer via diffusion transformers

Shengrong Gu, Ye Wang, Song Wu +4

Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively in…

cs.CV2026

Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation

Ying Tai, Nikai Du, Rui Xie +5

In this paper, we present TextCrafter, a Complex Visual Text Generation (CVTG) framework inspired by selective visual attention in cognitive science, and introduce the "Text Insula…