collaborators

5 papers

cs.LG2026

Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B

Jaeyeon Kim, Jewon Lee, Bo-Kyeong Kim

This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.5-4B on a resource-constrained NVIDIA A10G GPU. Our s…

cs.CV2026

Topology-Aware Layer Pruning for Large Vision-Language Models

Pengcheng Zheng, Chaoning Zhang, Ya Wen +10

Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding and reasoning, while recent extensions that incorporate visual inputs enable th…

cs.CL2026

Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking

Fachrina Dewi Puspitasari, Chaoning Zhang, Jiaquan Zhang +6

The demand for powerful instruction following and reasoning capability of large language models (LLMs) has promoted rapid development of retrieval-augmented generation (RAG). The R…

cs.CV2026

ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models

Jewon Lee, Wooksu Shin, Seungmin Yang +5

Efficient processing of high-resolution images is crucial for real-world vision-language applications. However, existing Large Vision-Language Models (LVLMs) incur substantial comp…

cs.CV2025

Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features

Jewon Lee, Ki-Ung Song, Seungmin Yang +6

Visual token reduction lowers inference costs caused by extensive image features in large vision-language models (LVLMs). Unlike relevant studies that prune tokens in self-attentio…