5 papers
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
Qing Yang, Pengcheng Huang, Xinze Li +6
Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and temporally dispersed across lengthy…
MARDoc: A Memory-Aware Refinement Agent Framework for Multimodal Long Document QA
Kaifeng Chen, Hongtao Liu, Qiyao Peng +4
Iterative retrieval-reasoning agents have recently shown promise for multimodal long-document question answering. However, most existing systems maintain a single growing context t…
PersonaAct: Simulating Short-Video Users with Personalized Agents for Counterfactual Filter Bubble Auditing
Shilong Zhao, Qinggang Yang, Zhiyi Yin +4
Short-video platforms rely on personalized recommendation, raising concerns about filter bubbles that narrow content exposure. Auditing such phenomena at scale is challenging becau…
Beyond Fixed Length: Bucket Pre-training is All You Need
Qing Yang, Qiyao Peng, Hongtao Liu +3
Large Language Models (LLMs) have demonstrated exceptional performance across various tasks, with pre-training stage serving as the cornerstone of their capabilities. However, the…
A Survey on LLM-powered Agents for Recommender Systems
Qiyao Peng, Hongtao Liu, Hua Huang +2
Recommender systems are essential components of many online platforms, yet traditional approaches still struggle with understanding complex user preferences and providing explainab…