collaborators

5 papers

cs.CV2026

Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models

Anmin Wang, Nan Zhang, Wei Tao +4

Vision-Language Models (VLMs) face significant computational challenges in video processing due to massive data redundancy, which creates prohibitively long token sequences. To add…

cs.CV2025

RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models

Junjie Li, Nan Zhang, Xiaoyang Qu +4

Object Navigation (ObjectNav) is a fundamental task in embodied artificial intelligence. Although significant progress has been made in semantic map construction and target directi…

cs.CL2025

MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs

Wei Tao, Xiaoyang Qu, Kai Lu +3

When applying pre-trained large language models (LLMs) to address anomaly detection tasks, the multivariate time series (MTS) modality of anomaly detection does not align with the…

cs.CV2025

VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection

Bin Zhang, Xiaoyang Qu, Guokuan Li +2

As object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot obj…

cs.CV2025

RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations

Bin Zhang, Jinggang Chen, Xiaoyang Qu +5

Enabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do no…