10 papers
Hierarchical Text-Guided Brain Tumor Segmentation via Sub-Region-Aware Prompts
Bahram Mohammadi, Ta Duc Huy, Afrouz Sheikholeslami +8
Brain tumor segmentation remains challenging because the three standard sub-regions, i.e., whole tumor (WT), tumor core (TC), and enhancing tumor (ET), often exhibit ambiguous visu…
Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification
Dexuan Ding, Ciyuan Peng, Endrowednes Kuantama +6
High-dimensional structural MRI (sMRI) images are widely used for Alzheimer's Disease (AD) diagnosis. Most existing methods for sMRI representation learning rely on 3D architecture…
Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos
Jianbo Ma, Hui Luo, Qi Chen +5
Multi-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos,…
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
Yunchuan Ma, Laiyun Qing, Guorong Li +4
Despite the significant progress of fully-supervised video captioning, zero-shot methods remain much less explored. In this paper, we propose a novel zero-shot video captioning fra…
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
Gaoxiang Cong, Liang Li, Jiadong Pan +5
Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief r…
FedDPG: An Adaptive Yet Efficient Prompt-tuning Approach in Federated Learning Settings
Ali Shakeri, Wei Emma Zhang, Amin Beheshti +3
Pre-trained Language Models (PLMs) have demonstrated impressive performance in various NLP tasks. However, traditional fine-tuning methods for leveraging PLMs for downstream tasks…