collaborators

6 papers

cs.CV2026

Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

Zhuo Song, Lian Xu, Runqing Jiang +4

Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods…

cs.CV2026

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

Lian Xu, Mohammed Bennamoun, Farid Boussaid +3

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily opera…

cs.CV2026

Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs

Wenjie Zhu, Yabin Zhang, Liang Xu +3

While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in…

cs.CV2026

ScalePredictor: Instance-aware Scale Learning for Accurate Quantization of Vision Transformers

Changjun Li, Runqing Jiang, Lian Xu +3

Vision Transformers have achieved remarkable success in many fields, yet their deployment on edge devices remains challenging due to their substantial computational demands. Post-T…

cs.CV2026

Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation

Yiheng Lyu, Lian Xu, Coen Arrow +3

Weakly supervised segmentation enables model training from plane-level labels. Existing methods often rely on 2D encoders, neglecting the volumetric nature of medical data. We prop…

cs.CV2026

Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection

A S M Sharifuzzaman Sagar, Mohammed Bennamoun, Farid Boussaid +4

In multimodal misinformation, deception usually arises not just from pixel-level manipulations in an image, but from the semantic and contextual claim jointly expressed by the imag…