366 citations · 952 across the 45 of their papers we have counts for
14 papers · 1 filter
Improving Dual-Encoder Training through Dynamic Indexes for Negative Mining
Nicholas Monath, Manzil Zaheer, Kelsey Allen +1
Dual encoder models are ubiquitous in modern classification and retrieval. Crucial for training such dual encoders is an accurate estimation of gradients from the partition functio…
Exact and Approximate Hierarchical Clustering Using A*
Craig S. Greenberg, Sebastian Macaluso, Nicholas Monath +6
Hierarchical clustering is a critical task in numerous domains. Many approaches are based on heuristics and the properties of the resulting clusterings are studied post hoc. Howeve…
Improving Local Identifiability in Probabilistic Box Embeddings
Shib Sankar Dasgupta, Michael Boratko, Dongxu Zhang +3
Geometric embeddings have recently received attention for their natural ability to represent transitive asymmetric relations via containment. Box embeddings, where objects are repr…
Scalable Hierarchical Agglomerative Clustering
Nicholas Monath, Avinava Dubey, Guru Guruganesh +9
The applicability of agglomerative clustering, for inferring both hierarchical and flat clustering, is limited by its scalability. Existing scalable hierarchical clustering methods…
Scalable Hierarchical Clustering with Tree Grafting
Nicholas Monath, Ari Kobren, Akshay Krishnamurthy +2
We introduce Grinch, a new algorithm for large-scale, non-greedy hierarchical clustering with general linkage functions that compute arbitrary similarity between two point sets. Th…
Optimal Transport-based Alignment of Learned Character Representations for String Similarity
Derek Tam, Nicholas Monath, Ari Kobren +3
String similarity models are vital for record linkage, entity resolution, and search. In this work, we present STANCE --a learned model for computing the similarity of two strings.…