Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding
Lin Chen, Bolin Ni, Qi Yang +5
Despite the remarkable capabilities of Multimodal Large Language Models (MLLMs), they still suffer from visual fading in long-context scenarios. Specifically, the attention to visu…
cs.CV2026
Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm
Sixing Li, Zhibin Gu, Ziqi Zhang +4
Image captioning for Early Childhood Education (ECE) is essential for automated activity understanding and educational assessment. However, existing methods face two key challenges…
cs.CV2024
LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
Ying Wang, Yanlai Yang, Mengye Ren
In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval. Lifelon…