collaborators

6 papers

cs.CV2026

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Deyi Zhu, Haoyu Fan, Yinan Zhu +4

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate traje…

cs.CV2026

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

Jingqi Tian, Haoji Zhang, Lin Chen +7

Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large lan…

cs.CV2026

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models

Yue Wu, Changyuan Wang, Zixuan Wang +2

Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution heterogeneity and disparate outli…

cs.CV2026

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

Shilin Ma, Chubin Zhang, Changyuan Wang +6

Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most…

cs.CV2026

VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution

Yixuan Zhu, Shilin Ma, Haolin Wang +6

Recent advancements in visual autoregressive models (VAR) have demonstrated their effectiveness in image generation, highlighting their potential for real-world image super-resolut…

cs.CV2025

FADE: Frequency-Aware Diffusion Model Factorization for Video Editing

Yixuan Zhu, Haolin Wang, Shilin Ma +4

Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional appr…