71 citations · 129 across the 20 of their papers we have counts for
6 papers · 1 filter
Scaling TabPFN: Sketching and Feature Selection for Tabular Prior-Data Fitted Networks
Benjamin Feuer, Chinmay Hegde, Niv Cohen
Tabular classification has traditionally relied on supervised algorithms, which estimate the parameters of a prediction model using its training data. Recently, Prior-Data Fitted N…
Exploring Dataset-Scale Indicators of Data Quality
Benjamin Feuer, Chinmay Hegde
Modern computer vision foundation models are trained on massive amounts of data, incurring large economic and environmental costs. Recent research has suggested that improving data…
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
Benjamin Feuer, Yurong Liu, Chinmay Hegde +1
Existing deep-learning approaches to semantic column type annotation (CTA) have important shortcomings: they rely on semantic types which are fixed at training time; require a larg…
Distributionally Robust Classification on a Data Budget
Benjamin Feuer, Ameya Joshi, Minh Pham +1
Real world uses of deep learning require predictable model behavior under distribution shifts. Models such as CLIP show emergent natural distributional robustness comparable to hum…
When Do Neural Nets Outperform Boosted Trees on Tabular Data?
Duncan McElfresh, Sujay Khandagale, Jonathan Valverde +6
Tabular data is one of the most commonly used types of data in machine learning. Despite recent advances in neural nets (NNs) for tabular data, there is still an active discussion…
LiT Tuned Models for Efficient Species Detection
Andre Nakkab, Benjamin Feuer, Chinmay Hegde
Recent advances in training vision-language models have demonstrated unprecedented robustness and transfer learning effectiveness; however, standard computer vision datasets are im…