3 papers
cs.AI2026
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models
Changming Xiao, Zhenliang Ni, Jinhui He +2
As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key…
cs.CL2026
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation
Xuan Du, Qiangyu Yan, Wenshuo Li +4
The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, ta…
cs.CV2025
Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild
Wanpeng Hu, Haodi Liu, Lin Chen +4
Complex visual reasoning remains a key challenge today. Typically, the challenge is tackled using methodologies such as Chain of Thought (COT) and visual instruction tuning. Howeve…