Publications (22)
SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
Yongzhi Li, Saining Zhang, Yibing Chen +3
Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fid…
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Junhao Chen, Mingjin Chen, Henghaofan Zhang +10
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Bohan Wang, Zhongqi Yue, Fengda Zhang +15
We completely discard the conventional spatial prior in image representation and introduce a novel discrete visual tokenizer: Self-consistency Tokenizer (Selftok). At its design co…
Feedforward 3D Editing Learns from Semantic-Part Transformation
Jiawei Weng, Saining Zhang, Zhenxin Diao +4
3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generative paradigms, 3D AI generati…
Controllable Radar Simulation with Waveform Parameter Embedding
Weiqing Xiao, Hao Huang, Chonghao Zhong +9
Autonomous driving simulators still lack high-fidelity radar, even though radar is critical for robust perception in adverse weather. A key obstacle is that raw radar point clouds…
Imagine Before You Draw: Visual Prompt Engineering for Image Generation
Liyu Jia, Fengda Zhang, Jiachun Pan +7
Incorporating visual semantic representations as an intermediate step before image generation can reduce the modeling difficulty between text and images, thereby improving generati…
ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
Zihao Huang, Tianqi Liu, Zhaoxi Chen +7
Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage…
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
Tianqi Liu, Zhaoxi Chen, Zihao Huang +8
Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighti…
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
Shaocong Xu, Songlin Wei, Qizhe Wei +12
Transparent objects remain notoriously hard for perception systems: refraction, reflection and transmission break the assumptions behind stereo, ToF and purely discriminative monoc…
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
Wei Chow, Jiachun Pan, Yongyuan Liang +10
Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primari…
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
Rongxu Cui, Zongzheng Zhang, Jingrui Pang +11
Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address th…
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
Zongzheng Zhang, Jingrui Pang, Zhuo Yang +22
Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper control or single-arm dextero…
Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting
Nan Wang, Yuantao Chen, Lixing Xiao +11
Neural rendering techniques, including NeRF and Gaussian Splatting (GS), rely on photometric consistency to produce high-quality reconstructions. However, in real-world scenarios,…
One Video, One World: Turning Monocular Video into Physical 4D Scenes
Junhao Chen, Boran Zhang, Mingjin Chen +7
We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconst…
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
Baijun Ye, Minghui Qin, Saining Zhang +7
Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy…
Bunraku: Turning a Single Illustration into an Editable Live2D Character
Junhao Chen, Jingjia Mao, Dayong Li +6
The paper introduces Bunraku, a system that automatically creates a complete Live2D character—including layered RGBA images, deformation meshes, and animation keyposes—from a singl…
GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects
Licheng Shen, Saining Zhang, Honghan Li +4
Reconstructing articulated objects is essential for building digital twins of interactive environments. However, prior methods typically decouple geometry and motion by first recon…
CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting
Haoran Xu, Saining Zhang, Peishuo Li +15
Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly c…
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Junhao Chen, Mingjin Chen, Jingjia Mao +12
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…
Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty
Saining Zhang, Baijun Ye, Xiaoxue Chen +5
Robust and realistic rendering for large-scale road scenes is essential in autonomous driving simulation. Recently, 3D Gaussian Splatting (3D-GS) has made groundbreaking progress i…
Engine-Native Editable 3D World Reconstruction with Objects and Lighting
Junhao Chen, Xinghao Chen, Henghaofan Zhang +8
Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at…
High-Resolution Boundary Detection for Medical Image Segmentation with Piece-Wise Two-Sample T-Test Augmented Loss
Yucong Lin, Jinhua Su, Yuhang Li +12
Deep learning methods have contributed substantially to the rapid advancement of medical image segmentation, the quality of which relies on the suitable design of loss functions. P…