Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity
MingZe Tang, Jubal Chandy Jacob
Recent Vision-Language Models (VLMs) enable zero-shot classification by aligning images and text in a shared space, a promising approach for data-scarce conditions. However, the in…
cs.CV2025
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
MingZe Tang, Madiha Kazi
This study explores human action recognition using a three-class subset of the COCO image corpus, benchmarking models from simple fully connected networks to transformer architectu…