5 papers
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
Jaeyeon Kim, Jewon Lee, Bo-Kyeong Kim
This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.5-4B on a resource-constrained NVIDIA A10G GPU. Our s…
Topology-Aware Layer Pruning for Large Vision-Language Models
Pengcheng Zheng, Chaoning Zhang, Ya Wen +10
Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding and reasoning, while recent extensions that incorporate visual inputs enable th…
Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking
Fachrina Dewi Puspitasari, Chaoning Zhang, Jiaquan Zhang +6
The demand for powerful instruction following and reasoning capability of large language models (LLMs) has promoted rapid development of retrieval-augmented generation (RAG). The R…
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
Jewon Lee, Wooksu Shin, Seungmin Yang +5
Efficient processing of high-resolution images is crucial for real-world vision-language applications. However, existing Large Vision-Language Models (LVLMs) incur substantial comp…
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
Jewon Lee, Ki-Ung Song, Seungmin Yang +6
Visual token reduction lowers inference costs caused by extensive image features in large vision-language models (LVLMs). Unlike relevant studies that prune tokens in self-attentio…