collaborators

6 papers

cs.CV2025

Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning

Chendong Wang, Donglin Bai, Yifan Yang +11

We present \emph{Video-in-the-Loop} (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first \emph{localizing} question-relevant interval(s) with a…

cs.RO2025

AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation

Xin Ding, Jianyu Wei, Yifan Yang +10

Vision Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning…

cs.DC2025

Scaling LLM Test-Time Compute with Mobile NPU on Smartphones

Zixu Hao, Jianyu Wei, Tuowei Wang +5

Deploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This p…

cs.CV2025

AVA: Towards Agentic Video Analytics with Vision Language Models

Yuxuan Yan, Shiqi Jiang, Ting Cao +5

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their ad…

cs.LG2025

Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash

Fucheng Jia, Zewen Wu, Shiqi Jiang +7

Large language models (LLMs) are increasingly being deployed on mobile devices, but the limited DRAM capacity constrains the deployable model size. This paper introduces ActiveFlow…

cs.CV2025

StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

Xin Ding, Hao Wu, Yifan Yang +4

With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce StreamMi…