works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CV2026

Semantically Calibrated Evidence Composition for CT Vision-Language Learning

Guoliang You, Haifan Gong, Xiaomeng Chu

Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typically emphasize either globa…

cs.CV2026

Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining

Guoliang You, Haifan Gong, Xiaomeng Chu

Volumetric CT vision-language pretraining learns 3D representations from scan-report pairs, but global and anatomy-aware objectives supervise only correspondence: they establish wh…

cs.CV2026

Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding

Guoliang You, Xiaomeng Chu

The paper introduces OCP-CT, a framework that aligns organ‑conditioned radiological pattern tokens between CT scans and radiology reports using a mixture‑of‑experts and contrastive…

cs.RO2026

Learning Surgical Robotic Manipulation with 3D Spatial Priors

Yu Sheng, Lidian Wang, Xiaomeng Chu +6

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the…

cs.RO2025

GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions

Xiaomeng Chu, Jiajun Deng, Guoliang You +4

Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities…

cs.CV2025

RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion

Xiaomeng Chu, Jiajun Deng, Guoliang You +3

We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception…