2 papers
cs.CV2026
Cloak of Invisibility: Real-Time Privacy-Preserving Volumetric Video Streaming
Hossein Khalili, Philip Do, Alexander Vilesov +3
Volumetric video streaming turns privacy into a 3D, multi-view problem. Unlike ordinary video, where sensitive content can often be redacted frame by frame, RGB-D volumetric pipeli…
cs.CV2025
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Shijie Zhou, Alexander Vilesov, Xuehai He +7
Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemp…