3 papers
cs.CL2026
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation
Xuan Du, Qiangyu Yan, Wenshuo Li +4
The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, ta…
cs.CV2025
Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild
Wanpeng Hu, Haodi Liu, Lin Chen +4
Complex visual reasoning remains a key challenge today. Typically, the challenge is tackled using methodologies such as Chain of Thought (COT) and visual instruction tuning. Howeve…
cs.CV2024
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
Changming Xiao, Qi Yang, Feng Zhou +1
Diffusion models have revolted the field of text-to-image generation recently. The unique way of fusing text and image information contributes to their remarkable capability of gen…