4 papers
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
Qitong Wang, Yijun Liang, Ming Li +2
Vision-Language Navigation (VLN) enables robots to follow natural-language instructions in visually grounded environments, serving as a key capability for embodied robotic systems.…
Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding
Qitong Wang, Fan Du, Pranav Maneriker +2
The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) tasks increasingly critical. We…
When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models
Qitong Wang, Haoran Dai, Haotian Zhang +2
While diffusion models have revolutionized visual content generation, their rapid adoption has underscored the critical need to investigate vulnerabilities, e.g., to backdoor attac…
United States Muon Collider Community White Paper for the European Strategy for Particle Physics Update
A. Abdelhamid, D. Acosta, P. Affleck +301
This document is being submitted to the 2024-2026 European Strategy for Particle Physics Update (ESPPU) process on behalf of the US Muon Collider community, with its preparation co…