4 papers
TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
Fang Li, Shihao Zou, Weixin Si +3
Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintaining their relational consisten…
VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis
Fang Li, Yang Gao, Shihao Zou +5
High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its ima…
QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction
Yinglong Li, Donghui Shen, Xiaoyu Zhang +5
While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer…
HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video
Yiming Jiang, Hanzhang Tu, Wenfeng Song +5
Uncalibrated volumetric video streaming for human reconstruction is essential for holographic communication and AR/VR, yet remains challenging due to the need for temporal consiste…