2 papers
cs.CV2026
A More Word-like Image Tokenization for MLLMs
Hyun Lee, Hyemin Jeong, Yejin Kim +4
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding…
cs.CV2026
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
Sumin Kim, Hyemin Jeong, Mingu Kang +3
The exponential growth of video content necessitates effective video summarization to efficiently extract key information from long videos. However, current approaches struggle to…