◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Vassil Tashev

3 papers hereh-index 343 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CY1
  • cs.LG1
  • cs.SE1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.SE2025

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

Sean McGregor, Victor Lu, Vassil Tashev +8

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by…

cs.CY2025

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Shaona Ghosh, Heather Frase, Adina Williams +99

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…

cs.LG2024

Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts

Jacob Haimes, Cenny Wenner, Kunvar Thaman +4

The training data for many Large Language Models (LLMs) is contaminated with test data. This means that public benchmarks used to assess LLMs are compromised, suggesting a performa…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.