6 papers · 1 filter
MEME: Multi-entity & Evolving Memory Evaluation
Seokwon Jung, Alexander Rubinstein, Arnas Uselis +2
LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. While prior benchmarks evaluate on…
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
Alexander Rubinstein, Benjamin Raible, Martin Gubri +1
Evaluating modern machine learning models has become prohibitively expensive. Benchmarks such as LMMs-Eval and HELM demand thousands of GPU hours per model. Costly evaluation reduc…
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
Luca Scimeca, Alexander Rubinstein, Damien Teney +2
Spurious correlations in the data, where multiple cues are predictive of the target labels, often lead to a phenomenon known as shortcut learning, where a model relies on erroneous…
Studying Large Language Model Behaviors Under Context-Memory Conflicts With Real Documents
Evgenii Kortukov, Alexander Rubinstein, Elisa Nguyen +1
Retrieval-augmented generation (RAG) mitigates many problems of fully parametric language models, such as temporal degradation, hallucinations, and lack of grounding. In RAG, the m…
Scalable Ensemble Diversification for OOD Generalization and Detection
Alexander Rubinstein, Luca Scimeca, Damien Teney +1
Training a diverse ensemble of models has several practical applications such as providing candidates for model selection with better out-of-distribution (OOD) generalization, and…
Do Deep Neural Network Solutions Form a Star Domain?
Ankit Sonthalia, Alexander Rubinstein, Ehsan Abbasnejad +1
It has recently been conjectured that neural network solution sets reachable via stochastic gradient descent (SGD) are convex, considering permutation invariances (Entezari et al.,…