575 citations · 1.4k across the 22 of their papers we have counts for
30 papers · 1 filter
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
Jingjie Ning, João Coelho, Yibo Kong +5
LLM-powered search agents are increasingly being used for multi-step information seeking tasks, yet the IR community lacks empirical understanding of how agentic search sessions un…
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
João Coelho, Jingjie Ning, Jingyuan He +8
Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries. However, most…
Dwell in the Beginning: How Language Models Embed Long Documents for Dense Retrieval
João Coelho, Bruno Martins, João Magalhães +2
This study investigates the existence of positional biases in Transformer-based models for text representation learning, particularly in the context of web document retrieval. We b…
Building Retrieval Systems for the ClueWeb22-B Corpus
Harshit Mehrotra, Jamie Callan, Zhen Fan
The ClueWeb22 dataset containing nearly 10 billion documents was released in 2022 to support academic and industry research. The goal of this project was to build retrieval baselin…
ClueWeb22: 10 Billion Web Documents with Visual and Semantic Information
Arnold Overwijk, Chenyan Xiong, Xiao Liu +2
ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information. Its design was influenced by the need for a high qu…
Tevatron: An Efficient and Flexible Toolkit for Dense Retrieval
Luyu Gao, Xueguang Ma, Jimmy Lin +1
Recent rapid advancements in deep pre-trained language models and the introductions of large datasets have powered research in embedding-based dense retrieval. While several good r…