3 papers
cs.CV2026
Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang
This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen objec…
cs.CV2026
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
Duy Tran Thanh, Thien-Phuc Doan, Long Nguyen-Vu +1
Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text…
cs.CV2026
Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines
Duy Tran Thanh
Production vision pipelines silently degrade on blurry input, wasting compute on downstream OCR, retrieval, and vision-language model (VLM) calls that cannot recover a usable outpu…