5 papers
LocAnyMed: Vision-Language Grounding for Multimodal Medical Images
Zihan Wang, Tong Liu, Zhiwei Wang +6
Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. H…
Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models
Sihan Ma, Siyuan Liang, Dacheng Tao
Generative 3D models are deployed in gaming, robotics, and immersive creation, making source attribution critical: given a 3D asset, can we identify whether and which generative mo…
JOintGS: Joint Optimization of Cameras, Bodies and 3D Gaussians for In-the-Wild Monocular Reconstruction
Zihan Lou, Jinlong Fan, Sihan Ma +2
Reconstructing high-fidelity animatable 3D human avatars from monocular RGB videos remains challenging, particularly in unconstrained in-the-wild scenarios where camera parameters…
ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification
Sihan Ma, Qiming Wu, Ruotong Jiang +1
The proliferation of digital news media necessitates robust methods for verifying content veracity, particularly regarding the consistency between visual and textual information. T…
End-to-End HOI Reconstruction Transformer with Graph-based Encoding
Zhenrong Wang, Qi Zheng, Sihan Ma +3
With the diversification of human-object interaction (HOI) applications and the success of capturing human meshes, HOI reconstruction has gained widespread attention. Existing main…