3 papers
cs.CV2026
SpecPL: Disentangling Spectral Granularity for Prompt Learning
Jingtao Zhou, Xirui Kang, Feiyang Huang +1
Existing prompt learning for VLMs exhibits a modality asymmetry, predominantly optimizing text tokens while still relying on frozen visual encoder as holistic extractor and neglect…
cs.CV2025
ViTOC: Vision Transformer and Object-aware Captioner
Feiyang Huang
This paper presents ViTOC (Vision Transformer and Object-aware Captioner), a novel vision-language model for image captioning that addresses the challenges of accuracy and diversit…
cs.CV2025
OPCap:Object-aware Prompting Captioning
Feiyang Huang
In the field of image captioning, the phenomenon where missing or nonexistent objects are used to explain an image is referred to as object bias (or hallucination). To mitigate thi…