activity
20242026
most citedEntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning

1 citations · 1 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2026

Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation

Tianrui Hui, Shaofei Huang, Qisong Han +6

Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method…

cs.CV2026

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

Yu Qi, Hongyu Li, Shaofei Huang +6

In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs…

cs.CV2026

Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets

Kaifeng Chen, Lechao Cheng, Jiyang Li +6

Dataset distillation (DD) condenses large corpora into compact, information-rich subsets for efficient training and reuse. However, under noisy supervision, DD risks condensing cor…

cs.CV2026

OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics

Jinjie Shen, Zheng Huang, Yuchen Zhang +7

Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-con…

cs.CV2026

CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition

Xu Wang, Shengeng Tang, Wan Jiang +3

Continuous Sign Language Recognition (CSLR) has achieved remarkable progress in recent years; however, most existing methods are developed under single-view settings and thus remai…

cs.CV2026

OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL

Jinjie Shen, Jing Wu, Yaxiong Wang +5

Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinform…