6 papers
Uncertainty-Aware World Model for Aerial Image-Goal Navigation
Deyi Zhu, Haoyu Fan, Yinan Zhu +4
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate traje…
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
Jingqi Tian, Haoji Zhang, Lin Chen +7
Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large lan…
MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models
Yue Wu, Changyuan Wang, Zixuan Wang +2
Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution heterogeneity and disparate outli…
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
Shilin Ma, Chubin Zhang, Changyuan Wang +6
Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most…
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
Yixuan Zhu, Shilin Ma, Haolin Wang +6
Recent advancements in visual autoregressive models (VAR) have demonstrated their effectiveness in image generation, highlighting their potential for real-world image super-resolut…
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
Yixuan Zhu, Haolin Wang, Shilin Ma +4
Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional appr…