output
20172026
most citedGrowth of Two-dimensional Compound Materials: Controllability, Material Quality, and Growth Mechanism

267 citations

Showing cs.CVShow all

19 papers · 1 filter

cs.CV2026

USCNet: Transformer-Based Multimodal Fusion with Segmentation Guidance for Urolithiasis Classification

Changmiao Wang, Songqi Zhang, Yongquan Zhang +9

Kidney stone disease ranks among the most prevalent conditions in urology, and understanding the composition of these stones is essential for creating personalized treatment plans…

cs.CV2026

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation

Zhuohong Chen, Zhengxian Wu, Zirui Liao +6

Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and i…

cs.CV2025★ 17 cited

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

Baining Zhao, Jianjie Fang, Zichao Dai +8

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a ben…

cs.CV2025★ 5 cited

SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition

Feng Lu, Tong Jin, Xiangyuan Lan +4

Recent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a…

cs.CV2024★ 9 cited

HisynSeg: Weakly-Supervised Histopathological Image Segmentation via Image-Mixing Synthesis and Consistency Regularization

Zijie Fang, Yifeng Wang, Peizhang Xie +2

Tissue semantic segmentation is one of the key tasks in computational pathology. To avoid the expensive and laborious acquisition of pixel-level annotations, a wide range of studie…

cs.CV2024★ 7 cited

EDTformer: An Efficient Decoder Transformer for Visual Place Recognition

Tong Jin, Feng Lu, Shuyu Hu +2

Visual place recognition (VPR) aims to determine the general geographical location of a query image by retrieving visually similar images from a large geo-tagged database. To obtai…