From the 1 of 7 linked papers with an AI index.
7 papers
ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
Xiao Luo, Mingyang Du, Xin Zhou +5
The paper introduces ROAD, a framework that transfers semantic and structural knowledge from discriminative 3D foundation models into diffusion transformers for 3D shape generation…
Stabilizing Temporal Inference Dynamics for Online Surgical Phase Recognition
Yang Liu, Ning Zhu, Jingjing Peng +4
Online Surgical Phase Recognition (SPR) models can reach high frame-wise accuracy, yet their predictions often lack temporal stability, fragmenting workflow understanding and reduc…
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
Xin Zhou, Dingkang Liang, Xiwu Chen +4
Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene gen…
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
Xin Zhou, Dingkang Liang, Sifan Tu +6
Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to inc…
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
Xin Zhou, Dingkang Liang, Kaijin Chen +7
Video generation models have demonstrated remarkable performance, yet their broader adoption remains constrained by slow inference speeds and substantial computational costs, prima…
DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation
Jiazhe Guo, Yikang Ding, Xiwu Chen +8
Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-sce…