6 papers
FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning
Zehao Li, Hongwei Yu, Hao Jiang +7
Multimodal large language models (MLLMs) have substantially advanced video misinformation detection through unified multimodal reasoning, but they often rely on fixed-depth inferen…
HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene
Jianing Chen, Zehao Li, Yujun Cai +7
Reconstructing dynamic 3D scenes from monocular videos remains a fundamental challenge in 3D vision. While 3D Gaussian Splatting (3DGS) achieves real-time rendering in static setti…
From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting
Jianing Chen, Zehao Li, Yujun Cai +5
Dynamic 3D reconstruction from monocular videos remains difficult due to the ambiguity inferring 3D motion from limited views and computational demands of modeling temporally varyi…
STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering
Zehao Li, Hao Jiang, Yujun Cai +7
Although dynamic scene reconstruction has long been a fundamental challenge in 3D vision, the recent emergence of 3D Gaussian Splatting (3DGS) offers a promising direction by enabl…
SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion Constraint
Zhenlong Yuan, Zhidong Yang, Yujun Cai +6
Recently, patch-deformation methods have exhibited significant effectiveness in multi-view stereo owing to the deformable and expandable patches in reconstructing textureless areas…
GradiSeg: Gradient-Guided Gaussian Segmentation with Enhanced 3D Boundary Precision
Zehao Li, Wenwei Han, Yujun Cai +5
While 3D Gaussian Splatting enables high-quality real-time rendering, existing Gaussian-based frameworks for 3D semantic segmentation still face significant challenges in boundary…