activity
20242026
collaborators
Showing 2025Show all

6 papers · 1 filter

cs.CV2025

Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos

Jianbo Ma, Hui Luo, Qi Chen +5

Multi-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos,…

cs.LG2025

FedDPG: An Adaptive Yet Efficient Prompt-tuning Approach in Federated Learning Settings

Ali Shakeri, Wei Emma Zhang, Amin Beheshti +3

Pre-trained Language Models (PLMs) have demonstrated impressive performance in various NLP tasks. However, traditional fine-tuning methods for leveraging PLMs for downstream tasks…

cs.MM2025

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

Gaoxiang Cong, Liang Li, Jiadong Pan +5

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief r…

cs.CV2025

SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting

Yiming Zhao, Guorong Li, Laiyun Qing +5

Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by…

cs.CV2025

ProgRoCC: A Progressive Approach to Rough Crowd Counting

Shengqin Jiang, Linfei Li, Haokui Zhang +6

As the number of individuals in a crowd grows, enumeration-based techniques become increasingly infeasible and their estimates increasingly unreliable. We propose instead an estima…

cs.CV2025

The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning

Mingkai Tian, Guorong Li, Yuankai Qi +4

Zero-shot video captioning requires that a model generate high-quality captions without human-annotated video-text pairs for training. State-of-the-art approaches to the problem le…