activity
20242026
collaborators

5 papers

cs.CV2026

AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding

Haozhe Qi, Kevin Qu, Mahdi Rad +3

Long video understanding remains challenging for Multi-modal Large Language Models (MLLMs) due to high memory costs and context-length limits. Prior approaches mitigate this by sco…

cs.CV2026

Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models

Kevin Qu, Haozhe Qi, Mihai Dusmanu +3

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware rea…

cs.CV2025

Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

Massimiliano Viola, Kevin Qu, Nando Metzger +4

Depth completion upgrades sparse depth measurements into dense depth maps guided by a conventional image. Existing methods for this highly ill-posed task operate in tightly constra…

cs.CV2025

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

Bingxin Ke, Kevin Qu, Tianfu Wang +5

The success of deep learning in computer vision over the past decade has hinged on large labeled datasets and strong pretrained models. In data-scarce settings, the quality of thes…

cs.AI2024

BIGCity: A Universal Spatiotemporal Model for Unified Trajectory and Traffic State Data Analysis

Xie Yu, Jingyuan Wang, Yifan Yang +2

Typical dynamic ST data includes trajectory data (representing individual-level mobility) and traffic state data (representing population-level mobility). Traditional studies often…