6 papers
Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning
Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee
In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image. Recent studies show th…
Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism
Gunho Jung, Jeong-Woo Park, Seon Bin Kim +1
Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction. In a training-free zero-shot s…
ACoRN: Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
Singon Kim, Gunho Jung, Seong-Whan Lee
Abstractive compression utilizes smaller langauge models to condense query-relevant context, reducing computational costs in retrieval-augmented generation (RAG). However,retrieved…
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
Gunho Jung, Heejo Kong, Seong-Whan Lee
Dynamic facial expression recognition (DFER) aims to identify emotional states by modeling the temporal changes in facial movements across video sequences. A key challenge in DFER…
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
Geon Park, Seon Bin Kim, Gunho Jung +1
With recent advancements in text-to-image (T2I) models, effectively generating multiple instances within a single image prompt has become a crucial challenge. Existing methods, whi…
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
Minsu Koh, Beom-Chul Park, Heejo Kong +1
Neural operators have emerged as promising frameworks for learning mappings governed by partial differential equations (PDEs), serving as data-driven alternatives to traditional nu…