6 papers
DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
Raffi Khatchadourian
A financial AI agent can repeat a decision while changing the tools, order, or recorded arguments and results used to reach it. Outcome-only evaluation misses this variation, even…
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
Raffi Khatchadourian
LLM agents struggle with regulatory audit replay: when asked to reproduce a flagged transaction decision with identical inputs, many deployments fail to return consistent results.…
Forecasting Future Language: Context Design for Mention Markets
Sumin Kim, Jihoon Kwon, Yoon Kim +9
Mention markets, a type of prediction market in which contracts resolve based on whether a specified keyword is mentioned during a future public event, require accurate probabilist…
An Empirical Framework for Evaluating Semantic Preservation Using Hugging Face
Nan Jia, Anita Raja, Raffi Khatchadourian
As machine learning (ML) becomes an integral part of high-autonomy systems, it is critical to ensure the trustworthiness of learning-enabled software systems (LESS). Yet, the nonde…
LLM Output Drift: Cross-Provider Validation & Mitigation for Financial Workflows
Raffi Khatchadourian, Rolando Franco
Financial institutions deploy Large Language Models (LLMs) for reconciliations, regulatory reporting, and client communications, but nondeterministic outputs (output drift) undermi…
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
Raffi Khatchadourian, Tatiana Castro Vélez, Mehdi Bagherzadeh +2
Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code -- s…