2 citations · 3 across the 17 of their papers we have counts for
17 papers · 1 filter
SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation
Junyin Zhang, Siyu Huang, Jianxiong Ye +4
Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are promising for this task because…
Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation
Pengyu Jie, Wanquan Liu, Rui He +5
White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewp…
ReUnit: Multi-Granularity Visual Unitization for Long Video Understanding
Biao Tang, Xu Chen, Shuxiang Gou +4
Long-video understanding is constrained by the limited visual input capacity of video multimodal large language models (Video-MLLMs). Existing methods mainly optimize which content…
WildGHand: Learning Anti-Perturbation Gaussian Hand Avatars from Monocular In-the-Wild Videos
Hanhui Li, Xuan Huang, Wanquan Liu +5
Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-w…
SOFTooth: Semantics-Enhanced Order-Aware Fusion for Tooth Instance Segmentation
Xiaolan Li, Wanquan Liu, Pengcheng Li +2
Three-dimensional (3D) tooth instance segmentation remains challenging due to crowded arches, ambiguous tooth-gingiva boundaries, missing teeth, and rare yet clinically important t…
Diffusion-Guided Mask-Consistent Paired Mixing for Endoscopic Image Segmentation
Pengyu Jie, Wanquan Liu, Rui He +3
Augmentation for dense prediction typically relies on either sample mixing or generative synthesis. Mixing improves robustness but misaligned masks yield soft label ambiguity. Diff…