2 papers
cs.LG2026
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
Surendra Pathak, Bo Han
Although Large Vision Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, their scalability and deployment are constrained by massive computatio…
cs.CV2026
ASAP: Attention-Shift-Aware Pruning for Efficient LVLM Inference
Surendra Pathak, Bo Han
While Large Vision-Language Models (LVLMs) demonstrate exceptional multi-modal capabilities, the quadratic computational cost of processing high-resolution visual tokens remains a…