◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ofir Arviv

4 papers hereh-index 10300 citations16 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG2

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2026

Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data

Ofir Arviv, Kristjan Greenewald, Yotam Perlitz +3

The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and test…

cs.CL2026

DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation

Eliya Habba, Ofir Arviv, Itay Itzhak +5

Recent work found that LLMs are sensitive to a wide range of arbitrary prompt dimensions, including the type of delimiters, answer enumerators, instruction wording, and more. This…

cs.CL2024

Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench

Yotam Perlitz, Ariel Gera, Ofir Arviv +5

Recent advancements in Language Models (LMs) have catalyzed the creation of multiple benchmarks, designed to assess these models' general capabilities. A crucial task, however, is…

cs.LG2024

Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications

Alon Halfon, Shai Gretz, Ofir Arviv +6

Fine-tuning Large Language Models (LLMs) is an effective method to enhance their performance on downstream tasks. However, choosing the appropriate setting of tuning hyperparameter…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.