Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
Qihang Ma, Shengyu Li, Jie Tang +5
Multi-modal keyphrase prediction (MMKP) aims to advance beyond text-only methods by incorporating multiple modalities of input information to produce a set of conclusive phrases. T…
cs.CV2024
SDPose: Tokenized Pose Estimation via Circulation-Guide Self-Distillation
Sichen Chen, Yingyi Zhang, Siming Huang +7
Recently, transformer-based methods have achieved state-of-the-art prediction quality on human pose estimation(HPE). Nonetheless, most of these top-performing transformer-based mod…