5 papers
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Adnan El Assadi, Niklas Muennighoff, Jinhyuk Lee
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embeddin…
MVEB: Massive Video Embedding Benchmark
Adnan El Assadi, Roman Solomatin, Isaac Chung +13
We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classificati…
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…
MAEB: Massive Audio Embedding Benchmark
Adnan El Assadi, Isaac Chung, Chenghao Xiao +15
We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasonin…
HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
Adnan El Assadi, Isaac Chung, Roman Solomatin +2
Comparing human and model performance offers a valuable perspective for understanding the strengths and limitations of embedding models, highlighting where they succeed and where t…