4 papers
YEAST: Yet Another Sequential Test
Alexey Kurennoy, Majed Dodin, Tural Gurbanov +1
Online evaluation of machine learning models is typically conducted through A/B experiments. Sequential statistical tests are valuable tools for analysing these experiments, as the…
Building a Scalable, Effective, and Steerable Search and Ranking Platform
Marjan Celikik, Jacek Wasilewski, Ana Peleteiro Ramallo +7
Modern e-commerce platforms offer vast product selections, making it difficult for customers to find items that they like and that are relevant to their current session intent. Thi…
Retrieve, Annotate, Evaluate, Repeat: Leveraging Multimodal LLMs for Large-Scale Product Retrieval Evaluation
Kasra Hosseini, Thomas Kober, Josip Krapac +3
Evaluating production-level retrieval systems at scale is a crucial yet challenging task due to the limited availability of a large pool of well-trained human annotators. Large Lan…
What should I wear to a party in a Greek taverna? Evaluation for Conversational Agents in the Fashion Domain
Antonis Maronikolakis, Ana Peleteiro Ramallo, Weiwei Cheng +1
Large language models (LLMs) are poised to revolutionize the domain of online fashion retail, enhancing customer experience and discovery of fashion online. LLM-powered conversatio…