Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Harshita Chopra, Kshitish Ghate, Aylin Caliskan +3
Large Language Model (LLM) agents are increasingly deployed in settings where they interact with a wide variety of people, including users who are unclear, impatient, or reluctant…
cs.AI2024
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
Steven A. Lehr, Aylin Caliskan, Suneragiri Liyanage +1
How good a research scientist is ChatGPT? We systematically probed the capabilities of GPT-3.5 and GPT-4 across four central components of the scientific process: as a Research Lib…