activity
20232026
most citedConventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.IRShow all

5 papers · 1 filter

cs.IR20251 cited

Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data

Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh +1

We investigate improving the retrieval effectiveness of embedding models through the lens of corpus-specific fine-tuning. Prior work has shown that fine-tuning with queries generat…

cs.IR2025

Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation

Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh +1

While the current state-of-the-art dense retrieval models exhibit strong out-of-domain generalization, they might fail to capture nuanced domain-specific knowledge. In principle, f…

cs.IR20251 cited

Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges

Manveer Singh Tamber, Jimmy Lin

This work considers a black-box threat model in which adversaries attempt to propagate arbitrary non-relevant content in search. We show that retrievers, rerankers, and LLM relevan…

cs.IR2024

Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models

Manveer Singh Tamber, Jasper Xian, Jimmy Lin

Embedding models that generate dense vector representations of text are widely used and hold significant commercial value. Companies such as OpenAI and Cohere offer proprietary emb…

cs.IR2023

Scaling Down, LiTting Up: Efficient Zero-Shot Listwise Reranking with Seq2seq Encoder-Decoder Models

Manveer Singh Tamber, Ronak Pradeep, Jimmy Lin

Recent work in zero-shot listwise reranking using LLMs has achieved state-of-the-art results. However, these methods are not without drawbacks. The proposed methods rely on large L…