Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
Bingjun Luo, Tony Wang, Chaoqi Chen +1
Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual tokens required. To improve efficie…
cs.AI2026
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding
Bingjun Luo, Tony Wang, Hanqi Chen +1
Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficiently compressing visual tokens wh…