Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
Chuck Arvin
This study examines how user-provided suggestions affect Large Language Models (LLMs) in a simulated educational context, where sycophancy poses significant risks. Testing five dif…
cs.CL2025
Identifying Legal Holdings with LLMs: A Systematic Study of Performance, Scale, and Memorization
Chuck Arvin
As large language models (LLMs) continue to advance in capabilities, it is essential to assess how they perform on established benchmarks. In this study, we present a suite of expe…