5 papers
Spatio-Temporal LLM: Reasoning about Environments and Actions
Haozhen Zheng, Beitong Tian, Mingyuan Wu +3
Despite significant recent progress of Multimodal Large Language Models (MLLMs), current MLLMs are challenged by "spatio-temporal" prompts, i.e., prompts that refer to 1) the entir…
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
Mingyuan Wu, Jize Jiang, Haozhen Zheng +8
Vision Language Models (VLMs) have achieved remarkable success in a wide range of vision applications of increasing complexity and scales, yet choosing the right VLM model size inv…
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
Beitong Tian, Lingzhi Zhao, Bo Chen +5
Underwater activities like scuba diving enable millions annually to explore marine environments for recreation and scientific research. Maintaining situational awareness and effect…
TraceNet: Segment one thing efficiently
Mingyuan Wu, Zichuan Liu, Haozhen Zheng +4
Efficient single instance segmentation is essential for unlocking features in the mobile imaging applications, such as capture or editing. Existing on-the-fly mobile imaging applic…
AquaScope: Reliable Underwater Image Transmission on Mobile Devices
Beitong Tian, Lingzhi Zhao, Bo Chen +5
Underwater communication is essential for both recreational and scientific activities, such as scuba diving. However, existing methods remain highly constrained by environmental ch…