7 papers
MVEB: Massive Video Embedding Benchmark
Adnan El Assadi, Roman Solomatin, Isaac Chung +13
We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classificati…
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs
Zacharie Bugaud
We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5 models (12B--70B) in 4,200 int…
CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms
Zacharie Bugaud
We introduce CheeseBench, a benchmark that evaluates large language models (LLMs) on nine classical behavioral neuroscience paradigms (Morris water maze, Barnes maze, T-maze, radia…
Cortex-Inspired Continual Learning: Unsupervised Instantiation and Recovery of Functional Task Networks
Kevin McKee, Thomas Hazy, Yicong Zheng +2
Block-sequential continual learning demands that a single model both protect prior solutions from catastrophic forgetting and efficiently infer at inference time which prior soluti…
Multi-RF Fusion with Multi-GNN Blending for Molecular Property Prediction
Zacharie Bugaud
Multi-RF Fusion achieves a test ROC-AUC of 0.8476 +/- 0.0002 on ogbg-molhiv (10 seeds), placing #1 on the OGB leaderboard ahead of HyperFusion (0.8475 +/- 0.0003). The core of the…
Hidden Clones: Exposing and Fixing Family Bias in Vision-Language Model Ensembles
Zacharie Bugaud
Ensembling Vision-Language Models (VLMs) from different providers maximizes benchmark accuracy, yet models from the same architectural family share correlated errors that standard…