8 papers
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
Jinlong Yang, Wenhao Zhang, Kuanwei Lin +1
Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example rega…
AnchorMem: Anchored Facts with Associative Contexts for Building Memory in Large Language Models
Zhanyu Shen, Sijie Cheng, Zhicheng Guo +3
While large language models have achieved remarkable performance in complex tasks, they still need a memory system to utilize historical experience in long-term interactions. Exist…
Evaluating Memory Capability in Continuous Lifelog Scenario
Jianjie Zheng, Zhichen Liu, Zhanyu Shen +6
Nowadays, wearable devices can continuously lifelog ambient conversations, creating substantial opportunities for memory systems. However, existing benchmarks primarily focus on on…
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
Junlong Li, Huaiyuan Xu, Sijie Cheng +4
Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is intro…
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
Hongbo Jin, Qingyuan Wang, Wenhao Zhang +2
Ultra long video understanding remains an open challenge, as existing vision language models (VLMs) falter on such content due to limited context length and inefficient long term m…
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
Peng Lai, Jianjie Zheng, Sijie Cheng +4
The growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as "LLM-as-a-judge". However, improving its alignment…