2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2026
Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback
Mohammad Al Ridhawi, Mahtab Haj Ali, Hussein Al Osman
Agentic artificial intelligence systems produce outputs through sequences of interdependent autonomous decisions, yet standard evaluation assesses outputs alone and cannot diagnose…
cs.CL2023★ 2 cited
ChatGPT for Suicide Risk Assessment on Social Media: Quantitative Evaluation of Model Performance, Potentials and Limitations
Hamideh Ghanadian, Isar Nejadgholi, Hussein Al Osman
This paper presents a novel framework for quantitatively evaluating the interactive ChatGPT model in the context of suicidality assessment from social media posts, utilizing the Un…