2 papers
cs.CV2025
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
Zeyi Huang, Utkarsh Ojha, Yuyang Ji +2
When a human undertakes a test, their responses likely follow a pattern: if they answered an easy question incorrectly, they would likely answer a more difficult one…
cs.CV2025
Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs
Zeyi Huang, Yuyang Ji, Xiaofang Wang +11
Long-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited cont…