From the 1 of 8 linked papers with an AI index.
8 papers
Semantically Calibrated Evidence Composition for CT Vision-Language Learning
Guoliang You, Haifan Gong, Xiaomeng Chu
Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typically emphasize either globa…
Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining
Guoliang You, Haifan Gong, Xiaomeng Chu
Volumetric CT vision-language pretraining learns 3D representations from scan-report pairs, but global and anatomy-aware objectives supervise only correspondence: they establish wh…
Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding
Guoliang You, Xiaomeng Chu
The paper introduces OCP-CT, a framework that aligns organ‑conditioned radiological pattern tokens between CT scans and radiology reports using a mixture‑of‑experts and contrastive…
Learning Surgical Robotic Manipulation with 3D Spatial Priors
Yu Sheng, Lidian Wang, Xiaomeng Chu +6
Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the…
GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions
Xiaomeng Chu, Jiajun Deng, Guoliang You +4
Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities…
RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion
Xiaomeng Chu, Jiajun Deng, Guoliang You +3
We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception…