3 papers
cs.CV2026
CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding
Zhenyu Yi, Qiang Hu, Zhenhao Li +3
Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that bu…
cs.CV2026
Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies
Zhenyu Yi, Zhiyun Song, Yusong Sun +6
Automated diagnosis of 3D brain CT scans is essential for critical care, yet it remains challenging due to the heavy reliance on manual annotations and the limited semantic underst…
cs.CV2026
LDGuid: A Framework for Robust Change Detection via Latent Difference Guidance
Jiaxuan Zhao, Ali Bereyhi
Modern deep learning models for change detection (CD) often struggle to explicitly represent task-relevant semantic differences. This paper proposes the Latent Difference Guidance…