From the 1 of 9 linked papers with an AI index.
9 papers
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
Mingyuan Wu, Jingcheng Yang, Shengyi Qian +11
The paper introduces SVR-R1, a reinforcement learning framework that lets a multimodal model generate an answer and then self‑verify it with a binary verdict, allowing a second‑cha…
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
Mingyuan Wu, Jingcheng Yang, Jize Jiang +6
Reinforcement Learning Finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, self-correction, a…
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Mingyuan Wu, Meitang Li, Jingcheng Yang +6
Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely…
Evaluating Spatio-Temporal Forecasting Trade-offs Between Graph Neural Networks and Foundation Models
Ragini Gupta, Naman Raina, Bo Chen +5
Modern IoT deployments for environmental sensing produce high volume spatiotemporal data to support downstream tasks such as forecasting, typically powered by machine learning mode…
Spatio-Temporal LLM: Reasoning about Environments and Actions
Haozhen Zheng, Beitong Tian, Mingyuan Wu +3
Despite significant recent progress of Multimodal Large Language Models (MLLMs), current MLLMs are challenged by "spatio-temporal" prompts, i.e., prompts that refer to 1) the entir…
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
Mingyuan Wu, Jize Jiang, Haozhen Zheng +8
Vision Language Models (VLMs) have achieved remarkable success in a wide range of vision applications of increasing complexity and scales, yet choosing the right VLM model size inv…