collaborators

6 papers

cs.CV2025

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

Muyao Yuan, Yuanhong Zhang, Weizhan Zhang +4

Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that…

cs.CV2025

GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models

Yushuo Zheng, Jiangyong Ying, Huiyu Duan +5

Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and…

cs.CV2025

Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition

Zefeng Qian, Xincheng Yao, Yifei Huang +3

Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recen…

cs.CV2025

RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment

Jianing Jin, Jiangyong Ying, Huiyu Duan +6

As camera-equipped robotic platforms become increasingly integrated into daily life, robotic-generated videos have begun to appear on streaming media platforms, enabling us to envi…

cs.CV2025

InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic Perspective

Yuanhong Zhang, Muyao Yuan, Weizhan Zhang +4

The Segment Anything Model (SAM), a vision foundation model, exhibits impressive zero-shot capabilities in general tasks but struggles in specialized domains. Parameter-efficient f…

cs.CV2025

Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition

Zefeng Qian, Chongyang Zhang, Yifei Huang +2

Few-shot Action Recognition (FSAR) constitutes a crucial challenge in computer vision, entailing the recognition of actions from a limited set of examples. Recent approaches mainly…