Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
OmniKVQuant: KV Cache Quantization for Omni-LLMs
Suho Yoo, Hyunjong Ok, Jongmin Choi +2
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-onl…
cs.CV2026
Out of Sight, Still in Mind: Token Compression for Omni-LLMs
Suho Yoo, Youngjoon Jang, Hyebin Cho +1
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but…
cs.CV2026
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
Suho Yoo, Youngjoon Jang, Joon Son Chung
The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process vi…