Showing 2024 · cs.CVShow all
3 papers · 2 filters
cs.CV2024
DistinctAD: Distinctive Audio Description Generation in Contexts
Bo Fang, Wenhao Wu, Qiangqiang Wu +2
Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automa…
cs.CV2024
Learning Tracking Representations from Single Point Annotations
Qiangqiang Wu, Antoni B. Chan
Existing deep trackers are typically trained with largescale video frames with annotated bounding boxes. However, these bounding boxes are expensive and time-consuming to annotate,…
cs.CV2024
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
Jia Wan, Qiangqiang Wu, Wei Lin +1
The existing crowd counting models require extensive training data, which is time-consuming to annotate. To tackle this issue, we propose a simple yet effective crowd counting meth…