6 papers · 1 filter
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
Yiling Gao, Hongchen Wei, Zhenzhong Chen
In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To a…
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
Zhihan Zhang, Xiang Pan, Hongchen Wei +1
Structural pruning techniques are essential for deploying multimodal large language models (MLLMs) across various hardware platforms, from edge devices to cloud servers. However, c…
RSFAKE-1M: A Large-Scale Dataset for Detecting Diffusion-Generated Remote Sensing Forgeries
Zhihong Tan, Jiayi Wang, Huiying Shi +3
Detecting forged remote sensing images is becoming increasingly critical, as such imagery plays a vital role in environmental monitoring, urban planning, and national security. Whi…
Training-Free Reasoning and Reflection in MLLMs
Hongchen Wei, Zhenzhong Chen
Recent advances in Reasoning LLMs (e.g., DeepSeek-R1 and OpenAI-o1) have showcased impressive reasoning capabilities via reinforcement learning. However, extending these capabiliti…
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
Hongchen Wei, Zhihong Tan, Yaosi Hu +2
Large Multimodal Models (LMMs) have demonstrated exceptional performance in video captioning tasks, particularly for short videos. However, as the length of the video increases, ge…
Visual Context Window Extension: A New Perspective for Long Video Understanding
Hongchen Wei, Zhenzhong Chen
Large Multimodal Models (LMMs) have demonstrated impressive performance in short video understanding tasks but face great challenges when applied to long video understanding. In co…