event-aware visual allocation 1gaussian mixture modeling 1keyframe selection 1long video understanding 1visual token budgeting 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +3
The paper introduces GMM-EVA, a training-free framework that uses Gaussian Mixture Models to detect event-level structures in long videos and allocate visual tokens by selecting on…
cs.CV2026
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Meituan LongCat Team, Bin Xiao, Chao Wang +86
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…
cs.CV2025
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +5
Large Vision-Language Models (LVLMs) suffer from serious hallucination problems, where the model-generated responses are inconsistent with the visual inputs. Existing hallucination…