◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jungkoo Kang

3 papers hereh-index 14 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2
  • middle author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI1
  • cs.LG1
  • cs.SE1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains

Jung Min Kang

Benchmark evaluation across AI and safety-critical domains overwhelmingly relies on simple averaging. We demonstrate that this practice produces substantially misleading rankings w…

cs.SE2026

Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling

Benjamin Elder, Anupama Murthi, Jungkoo Kang +4

Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in reali…

cs.AI2025

Scaling LLM Planning: NL2FLOW for Parametric Problem Generation and Rigorous Evaluation

Jungkoo Kang

Robust workflow composition is critical for effective agent performance, yet progress in Large Language Model (LLM) planning and reasoning is hindered by a scarcity of scalable eva…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.