1 paper
Renjie Liang, Zijian Xu, Jinqian Pan +6
A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed befor…