2 papers
cs.AI2026
OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance
Yeo Jeong Park, Hyemi Jang, Minseo Choi +3
Omni-modal large language models have demonstrated remarkable potential in holistic multimodal understanding; however, the token explosion caused by high-resolution audio and video…
cs.LG2026
TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
Junhan Kim, Yeo Jeong Park, Seungwoo Son +4
The rapid growth of large language models (LLMs) has heightened the importance of post-training quantization (PTQ) for reducing memory and computation costs. Among PTQ methods, GPT…