collaborators

6 papers

cs.CV2026

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Sunqi Fan, Qingle Liu, Runqi Yin +2

Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Quest…

cs.CV2026

PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation

Xin-Sheng Chen, Jiayu Zhu, Pei-lin Li +3

Slides serve as a critical medium for conveying information in presentation-oriented scenarios such as academia, education, and business. Despite their importance, creating high-qu…

cs.IR2026

Recurrent Preference Memory for Efficient Long-Sequence Generative Recommendation

Yixiao Chen, Yuan Wang, Yue Liu +9

Generative recommendation (GenRec) models typically model user behavior via full attention, but scaling to lifelong sequences is hindered by prohibitive computational costs and noi…

cs.IR2026

GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation

Jun Zhang, Yi Li, Yue Liu +19

As an intelligent infrastructure connecting users with commercial content, advertising recommendation systems play a central role in information flow and value creation within the…

cs.CV2025

Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task

Sunqi Fan, Jiashuo Cui, Meng-Hao Guo +1

Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand, and reason about dynamic real…

cs.CV2025

Agentic Keyframe Search for Video Question Answering

Sunqi Fan, Meng-Hao Guo, Shuojin Yang

Video question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards ach…