3 papers
cs.IR2025
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
João Coelho, Jingjie Ning, Jingyuan He +8
Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries. However, most…
cs.IR2025
ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests
Jingyuan He, Jiongnan Liu, Vishan Vishesh Oberoi +7
Recommender systems are among the most impactful AI applications, interacting with billions of users every day, guiding them to relevant products, services, or information tailored…
cs.CV2024
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Xinyu Fang, Kangrui Mao, Haodong Duan +4
The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA be…