1 citations · 1 across the 6 of their papers we have counts for
8 papers
Deep in the Jungle: Towards Automating Chimpanzee Population Estimation
Tom Raynes, Otto Brookes, Timm Haucke +6
The estimation of abundance and density in unmarked populations of great apes relies on statistical frameworks that require animal-to-camera distance measurements. In practice, acq…
Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
Connor Kilrain, David Carlyn, Julia Chae +3
The rise of personalized generative models raises a central question: how should we evaluate identity preservation? Given a reference image (e.g., one's pet), we expect the generat…
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
Olawale Salaudeen, Haoran Zhang, Kumail Alhamoud +2
Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-…
DataS^3: Dataset Subset Selection for Specialization
Neha Hulkund, Alaa Maalouf, Levi Cai +15
In many real-world machine learning (ML) applications (e.g. detecting broken bones in x-ray images, detecting species in camera traps), in practice models need to perform well on s…
Pairwise Matching of Intermediate Representations for Fine-grained Explainability
Lauren Shrack, Timm Haucke, Antoine Salaün +2
The differences between images belonging to fine-grained categories are often subtle and highly localized, and existing explainability techniques for deep learning models are often…
Do Large Language Model Benchmarks Test Reliability?
Joshua Vendrow, Edward Vendrow, Sara Beery +1
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' g…