3 papers
cs.CV2026
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
V Team, Wenyi Hong, Wenmeng Yu +90
We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this…
cs.CV2025
WakeupUrban: Unsupervised Semantic Segmentation of Mid-20 century Urban Landscapes with Satellite Imagery
Tianxiang Hao, Lixian Zhang, Yingjia Zhang +4
Historical satellite imagery archive, such as Keyhole satellite data, offers rare insights into understanding early urban development and long-term transformation. However, severe…
cs.LG2025
SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
Xiangwen Zhuge, Xu Shen, Zeyu Wang +6
Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model…