2 papers
cs.CV2025
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
Yogesh Kumar
Vision Language Models (VLMs) struggle with long-form videos due to the quadratic complexity of attention mechanisms. We propose Language-Guided Temporal Token Pruning (LGTTP), whi…
cs.CV2025
VideoLLM Benchmarks and Evaluation: A Survey
Yogesh Kumar
The rapid development of Large Language Models (LLMs) has catalyzed significant advancements in video understanding technologies. This survey provides a comprehensive analysis of b…