3 papers
cs.CV2026
Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering
Fan Wei, Siru Zhong, Runmin Dong +3
Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods…
cs.CV2026
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
Zhaoyang Luo, Runmin Dong, Miao Yang +4
Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remov…
cs.CV2026
RS-Prune: Training-Free Data Pruning at High Ratios for Efficient Remote Sensing Diffusion Foundation Models
Fan Wei, Runmin Dong, Yushan Lai +8
Diffusion-based remote sensing (RS) generative foundation models are cruial for downstream tasks. However, these models rely on large amounts of globally representative data, which…