activity
20102024
most citedEfficiently Inducing Features of Conditional Random Fields

366 citations · 952 across the 45 of their papers we have counts for

collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG20231 cited

Improving Dual-Encoder Training through Dynamic Indexes for Negative Mining

Nicholas Monath, Manzil Zaheer, Kelsey Allen +1

Dual encoder models are ubiquitous in modern classification and retrieval. Crucial for training such dual encoders is an accurate estimation of gradients from the partition functio…

cs.LG20212 cited

Exact and Approximate Hierarchical Clustering Using A*

Craig S. Greenberg, Sebastian Macaluso, Nicholas Monath +6

Hierarchical clustering is a critical task in numerous domains. Many approaches are based on heuristics and the properties of the resulting clusterings are studied post hoc. Howeve…

cs.LG202022 cited

Improving Local Identifiability in Probabilistic Box Embeddings

Shib Sankar Dasgupta, Michael Boratko, Dongxu Zhang +3

Geometric embeddings have recently received attention for their natural ability to represent transitive asymmetric relations via containment. Box embeddings, where objects are repr…

cs.LG2020

Scalable Hierarchical Agglomerative Clustering

Nicholas Monath, Avinava Dubey, Guru Guruganesh +9

The applicability of agglomerative clustering, for inferring both hierarchical and flat clustering, is limited by its scalability. Existing scalable hierarchical clustering methods…

cs.LG201917 cited

Scalable Hierarchical Clustering with Tree Grafting

Nicholas Monath, Ari Kobren, Akshay Krishnamurthy +2

We introduce Grinch, a new algorithm for large-scale, non-greedy hierarchical clustering with general linkage functions that compute arbitrary similarity between two point sets. Th…

cs.LG20191 cited

Optimal Transport-based Alignment of Learned Character Representations for String Similarity

Derek Tam, Nicholas Monath, Ari Kobren +3

String similarity models are vital for record linkage, entity resolution, and search. In this work, we present STANCE --a learned model for computing the similarity of two strings.…