13 papers
Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation
Tianrui Hui, Shaofei Huang, Qisong Han +6
Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method…
Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
Yu Qi, Hongyu Li, Shaofei Huang +6
In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs…
Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets
Kaifeng Chen, Lechao Cheng, Jiyang Li +6
Dataset distillation (DD) condenses large corpora into compact, information-rich subsets for efficient training and reuse. However, under noisy supervision, DD risks condensing cor…
Multi-Scale Global-Instance Prompt Tuning for Continual Test-time Adaptation in Medical Image Segmentation
Lingrui Li, Yanfeng Zhou, Nan Pu +2
Distribution shift is a common challenge in medical images obtained from different clinical centers, significantly hindering the deployment of pre-trained semantic segmentation mod…
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
Qihang Wang, Yaxiong Wang, Lechao Cheng +1
This paper explores image editing under the joint control of text and drag interactions. While recent advances in text-driven and drag-driven editing have achieved remarkable progr…
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
Jinjie Shen, Yaxiong Wang, Lechao Cheng +2
The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progre…