8 papers · 1 filter
Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
Zanwei Zhou, Taoran Yi, Jiemin Fang +5
Flow-based 3D generation models typically require dozens of sampling steps during inference. Though few-step distillation methods, particularly Consistency Models (CMs), have achie…
How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study
Che Liu, Jiazhen Pan, Weixiang Shen +3
Vision-Language Models (VLMs) trained on web-scale corpora excel at natural image tasks and are increasingly repurposed for healthcare; however, their competence in medical tasks r…
Tackling View-Dependent Semantics in 3D Language Gaussian Splatting
Jiazhong Cen, Xudong Zhou, Jiemin Fang +5
Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabul…
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation
Zelin Peng, Zhengqin Xu, Zhilin Zeng +2
Open-vocabulary semantic segmentation seeks to label each pixel in an image with arbitrary text descriptions. Vision-language foundation models, especially CLIP, have recently emer…
ViTree: Single-path Neural Tree for Step-wise Interpretable Fine-grained Visual Categorization
Danning Lao, Qi Liu, Jiazi Bu +2
As computer vision continues to advance and finds widespread applications across various domains, the need for interpretability in deep learning models becomes paramount. Existing…
Segment Any 3D Gaussians
Jiazhong Cen, Jiemin Fang, Chen Yang +4
This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input,…