2 papers
cs.CV2026
Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs
Zeyi Huang, Yuyang Ji, Xiaofang Wang +11
Long-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited cont…
cs.CV2025
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
Zeyi Huang, Utkarsh Ojha, Yuyang Ji +2
When a human undertakes a test, their responses likely follow a pattern: if they answered an easy question incorrectly, they would likely answer a more difficult one…