3 papers
cs.CV2025
Bridging the gap to real-world language-grounded visual concept learning
Whie Jung, Semin Kim, Junee Kim +1
Human intelligence effortlessly interprets visual scenes along a rich spectrum of semantic dimensions. However, existing approaches to language-grounded visual concept learning are…
cs.LG2025
Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models
Semin Kim, Yeonwoo Cha, Jaehoon Yoo +1
We investigate a general approach for improving user prompts in text-to-image (T2I) diffusion models by finding prompts that maximize a reward function specified at test-time. Alth…
cs.CV2025
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
Yoorhim Cho, Hongyeob Kim, Semin Kim +3
Visuo-tactile perception aims to understand an object's tactile properties, such as texture, softness, and rigidity. However, the field remains underexplored because collecting tac…