papers

Publications (22)

cs.GR2025

SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation

Yongzhi Li, Saining Zhang, Yibing Chen +3

Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fid…

cs.CV2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

Junhao Chen, Mingjin Chen, Henghaofan Zhang +10

Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…

cs.CV2025

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Bohan Wang, Zhongqi Yue, Fengda Zhang +15

We completely discard the conventional spatial prior in image representation and introduce a novel discrete visual tokenizer: Self-consistency Tokenizer (Selftok). At its design co…

cs.CV2026

Feedforward 3D Editing Learns from Semantic-Part Transformation

Jiawei Weng, Saining Zhang, Zhenxin Diao +4

3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generative paradigms, 3D AI generati…

eess.SP2026

Controllable Radar Simulation with Waveform Parameter Embedding

Weiqing Xiao, Hao Huang, Chonghao Zhong +9

Autonomous driving simulators still lack high-fidelity radar, even though radar is critical for robust perception in adverse weather. A key obstacle is that raw radar point clouds…

cs.CV2026

Imagine Before You Draw: Visual Prompt Engineering for Image Generation

Liyu Jia, Fengda Zhang, Jiachun Pan +7

Incorporating visual semantic representations as an intermediate step before image generation can reduce the modeling difficulty between text and images, thereby improving generati…

cs.CV2026

ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

Zihao Huang, Tianqi Liu, Zhaoxi Chen +7

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage…

cs.CV2025

Light-X: Generative 4D Video Rendering with Camera and Illumination Control

Tianqi Liu, Zhaoxi Chen, Zihao Huang +8

Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighti…

cs.CV2025

Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation

Shaocong Xu, Songlin Wei, Qizhe Wei +12

Transparent objects remain notoriously hard for perception systems: refraction, reflection and transmission break the assumptions behind stereo, ToF and purely discriminative monoc…

cs.CV2025

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

Wei Chow, Jiachun Pan, Yongyuan Liang +10

Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primari…

cs.RO2026

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

Rongxu Cui, Zongzheng Zhang, Jingrui Pang +11

Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address th…

cs.RO2026

Dexora: Open-source VLA for High-DoF Bimanual Dexterity

Zongzheng Zhang, Jingrui Pang, Zhuo Yang +22

Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper control or single-arm dextero…

cs.CV2026

Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting

Nan Wang, Yuantao Chen, Lixing Xiao +11

Neural rendering techniques, including NeRF and Gaussian Splatting (GS), rely on photometric consistency to produce high-quality reconstructions. However, in real-world scenarios,…

cs.CV2026

One Video, One World: Turning Monocular Video into Physical 4D Scenes

Junhao Chen, Boran Zhang, Mingjin Chen +7

We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconst…

cs.CV2025

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

Baijun Ye, Minghui Qin, Saining Zhang +7

Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy…

cs.CV2026

Bunraku: Turning a Single Illustration into an Editable Live2D Character

Junhao Chen, Jingjia Mao, Dayong Li +6

The paper introduces Bunraku, a system that automatically creates a complete Live2D character—including layered RGBA images, deformation meshes, and animation keyposes—from a singl…

#character animation#image decomposition#mesh generation#diffusion models
cs.CV2026

GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects

Licheng Shen, Saining Zhang, Honghan Li +4

Reconstructing articulated objects is essential for building digital twins of interactive environments. However, prior methods typically decouple geometry and motion by first recon…

cs.CV2025

CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting

Haoran Xu, Saining Zhang, Peishuo Li +15

Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly c…

cs.SD2026

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Junhao Chen, Mingjin Chen, Jingjia Mao +12

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…

cs.CV2024

Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty

Saining Zhang, Baijun Ye, Xiaoxue Chen +5

Robust and realistic rendering for large-scale road scenes is essential in autonomous driving simulation. Recently, 3D Gaussian Splatting (3D-GS) has made groundbreaking progress i…

cs.CV2026

Engine-Native Editable 3D World Reconstruction with Objects and Lighting

Junhao Chen, Xinghao Chen, Henghaofan Zhang +8

Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at…

eess.IV2022

High-Resolution Boundary Detection for Medical Image Segmentation with Piece-Wise Two-Sample T-Test Augmented Loss

Yucong Lin, Jinhua Su, Yuhang Li +12

Deep learning methods have contributed substantially to the rapid advancement of medical image segmentation, the quality of which relies on the suitable design of loss functions. P…