2 papers
cs.CV2026
When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation
Renjie Liang, Zijian Xu
Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language m…
cs.CV2026
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
Renjie Liang, Zijian Xu, Jinqian Pan +6
A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed befor…