2 papers
cs.CV2026
PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models
Jihyung Ko, Eunji Jung, Hyeongsub Kim +4
Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. E…
cs.AI2026
On the Collapse of Generative Paths: A Criterion and Correction for Diffusion Steering
Ziseok Lee, Minyeong Hwang, Wooyeol Lee +6
Inference-time steering adapts pretrained diffusion and flow models to new tasks without retraining, often utilizing ratio-of-densities constructions that reweight time-indexed mar…