2 citations · 2 across the 1 of their papers we have counts for
4 papers
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
Juyeon Yoon, Somin Kim, Robert Feldt +1
Software increasingly relies on the emergent capabilities of Large Language Models (LLMs), from natural language understanding to program analysis and generation. Yet testing them…
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
Junho Myung, Yeon Su Park, Sunwoo Kim +2
Evaluating the performance and biases of large language models (LLMs) through role-playing scenarios is becoming increasingly common, as LLMs often exhibit biased behaviors in thes…
Capturing Semantic Flow of ML-based Systems
Shin Yoo, Robert Feldt, Somin Kim +1
ML-based systems are software systems that incorporates machine learning components such as Deep Neural Networks (DNNs) or Large Language Models (LLMs). While such systems enable a…
DANDI: Diffusion as Normative Distribution for Deep Neural Network Input
Somin Kim, Shin Yoo
Surprise Adequacy (SA) has been widely studied as a test adequacy metric that can effectively guide software engineers towards inputs that are more likely to reveal unexpected beha…