6 papers
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Sunqi Fan, Qingle Liu, Runqi Yin +2
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Quest…
PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
Xin-Sheng Chen, Jiayu Zhu, Pei-lin Li +3
Slides serve as a critical medium for conveying information in presentation-oriented scenarios such as academia, education, and business. Despite their importance, creating high-qu…
Recurrent Preference Memory for Efficient Long-Sequence Generative Recommendation
Yixiao Chen, Yuan Wang, Yue Liu +9
Generative recommendation (GenRec) models typically model user behavior via full attention, but scaling to lifelong sequences is hindered by prohibitive computational costs and noi…
GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation
Jun Zhang, Yi Li, Yue Liu +19
As an intelligent infrastructure connecting users with commercial content, advertising recommendation systems play a central role in information flow and value creation within the…
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
Sunqi Fan, Jiashuo Cui, Meng-Hao Guo +1
Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand, and reason about dynamic real…
Agentic Keyframe Search for Video Question Answering
Sunqi Fan, Meng-Hao Guo, Shuojin Yang
Video question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards ach…