6 papers
PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models
Jihyung Ko, Eunji Jung, Hyeongsub Kim +4
Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. E…
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3
Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing semantics among similar objects.…
On the Collapse of Generative Paths: A Criterion and Correction for Diffusion Steering
Ziseok Lee, Minyeong Hwang, Wooyeol Lee +6
Inference-time steering adapts pretrained diffusion and flow models to new tasks without retraining, often utilizing ratio-of-densities constructions that reweight time-indexed mar…
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
Sanghyun Jo, Ziseok Lee, Wooyeol Lee +3
High-quality instance and panoptic segmentation has traditionally relied on dense instance-level annotations such as masks, boxes, or points, which are costly, inconsistent, and di…
Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
Joowon Kim, Ziseok Lee, Donghyeon Cho +4
Despite recent advances in diffusion models, achieving reliable image generation and editing remains challenging due to the inherent diversity induced by stochastic noise in the sa…
HybridLinker: Topology-Guided Posterior Sampling for Enhanced Diversity and Validity in 3D Molecular Linker Generation
Minyeong Hwang, Ziseok Lee, Kwang-Soo Kim +2
Linker generation is critical in drug discovery applications such as lead optimization and PROTAC design, where molecular fragments are assembled into diverse drug candidates via m…