5 papers
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
Paribesh Regmi, Qingshuang Chen, Chi Zhang +3
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deploy…
Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
Taeho Kang, Nairan Zhang, Yelin Kim +2
Accurate yaw estimation is a bottleneck in cross-view localization between ground view and Bird's Eye View (BEV). Existing methods couple yaw with translation and rely on height or…
AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization
Mohammad Omama, Gabriele Berton, Eric Foxlin +1
Precise and real-time visual localization is critical for applications like AR/VR and robotics, especially on resource-constrained edge devices such as smart glasses, where battery…
studentSplat: Your Student Model Learns Single-view 3D Gaussian Splatting
Yimu Pan, Hongda Mao, Qingshuang Chen +1
Recent advance in feed-forward 3D Gaussian splatting has enable remarkable multi-view 3D scene reconstruction or single-view 3D object reconstruction but single-view 3D scene recon…
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
Sungjune Park, Hongda Mao, Qingshuang Chen +2
As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. Howe…