267 citations
- Tsinghua UniversityCN100 papers
- University Town of ShenzhenCN21 papers
- Chinese Academy of SciencesCN15 papers
- Peng Cheng LaboratoryCN10 papers
- Peking UniversityCN8 papers
- Tencent (China)CN8 papers
- Tsinghua Shenzhen International Graduate SchoolCN8 papers
- Southern University of Science and TechnologyCN7 papers
- Nanyang Technological UniversitySG4 papers
- Shandong UniversityCN4 papers
- University of ManchesterGB4 papers
- École Polytechnique Fédérale de LausanneCH3 papers
19 papers · 1 filter
USCNet: Transformer-Based Multimodal Fusion with Segmentation Guidance for Urolithiasis Classification
Changmiao Wang, Songqi Zhang, Yongquan Zhang +9
Kidney stone disease ranks among the most prevalent conditions in urology, and understanding the composition of these stones is essential for creating personalized treatment plans…
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation
Zhuohong Chen, Zhengxian Wu, Zirui Liao +6
Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and i…
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
Baining Zhao, Jianjie Fang, Zichao Dai +8
Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a ben…
SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition
Feng Lu, Tong Jin, Xiangyuan Lan +4
Recent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a…
HisynSeg: Weakly-Supervised Histopathological Image Segmentation via Image-Mixing Synthesis and Consistency Regularization
Zijie Fang, Yifeng Wang, Peizhang Xie +2
Tissue semantic segmentation is one of the key tasks in computational pathology. To avoid the expensive and laborious acquisition of pixel-level annotations, a wide range of studie…
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
Tong Jin, Feng Lu, Shuyu Hu +2
Visual place recognition (VPR) aims to determine the general geographical location of a query image by retrieving visually similar images from a large geo-tagged database. To obtai…