8 papers
Evaluating Spatio-Temporal Forecasting Trade-offs Between Graph Neural Networks and Foundation Models
Ragini Gupta, Naman Raina, Bo Chen +5
Modern IoT deployments for environmental sensing produce high volume spatiotemporal data to support downstream tasks such as forecasting, typically powered by machine learning mode…
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
Beitong Tian, Lingzhi Zhao, Bo Chen +5
Underwater activities like scuba diving enable millions annually to explore marine environments for recreation and scientific research. Maintaining situational awareness and effect…
Spatio-Temporal LLM: Reasoning about Environments and Actions
Haozhen Zheng, Beitong Tian, Mingyuan Wu +3
Despite significant recent progress of Multimodal Large Language Models (MLLMs), current MLLMs are challenged by "spatio-temporal" prompts, i.e., prompts that refer to 1) the entir…
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Mingyuan Wu, Meitang Li, Jingcheng Yang +6
Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely…
Performance Characterization of Containers in Edge Computing
Ragini Gupta, Klara Nahrstedt
Edge computing addresses critical limitations of cloud computing such as high latency and network congestion by decentralizing processing from cloud to the edge. However, the need…
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
Mingyuan Wu, Jingcheng Yang, Jize Jiang +6
Reinforcement Learning Finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, self-correction, a…