2 papers
cs.CV2025
Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild
Wanpeng Hu, Haodi Liu, Lin Chen +4
Complex visual reasoning remains a key challenge today. Typically, the challenge is tackled using methodologies such as Chain of Thought (COT) and visual instruction tuning. Howeve…
cs.CV2024
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
Changming Xiao, Qi Yang, Feng Zhou +1
Diffusion models have revolted the field of text-to-image generation recently. The unique way of fusing text and image information contributes to their remarkable capability of gen…