15 papers
DiG-bench: Discovery in Games
Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few…
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
Hoyoung Lee, Suhwan Park, Seunghan Lee +15
Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large language models (LLMs) compress financial sourc…
Evaluating LLMs in Finance Requires Explicit Bias Consideration
Yaxuan Kong, Hoyoung Lee, Yoontae Hwang +7
Large Language Models (LLMs) are increasingly integrated into financial workflows, but evaluation practice has not kept up. Finance-specific biases can inflate performance, contami…
Deep Reinforcement Learning for Optimum Order Execution: Mitigating Risk and Maximizing Returns
Khabbab Zakaria, Jayapaulraj Jerinsh, Andreas Maier +3
Optimal Order Execution is a well-established problem in finance that pertains to the flawless execution of a trade (buy or sell) for a given volume within a specified time frame.…
Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI
Samarth Sarin, Lovepreet Singh, Bhaskarjit Sarmah +1
Agentic memory is emerging as a key enabler for large language models (LLM) to maintain continuity, personalization, and long-term context in extended user interactions, critical c…
A Unified AI System For Data Quality Control and DataOps Management in Regulated Environments
Devender Saini, Bhavika Jain, Nitish Ujjwal +3
In regulated domains such as finance, the integrity and governance of data pipelines are critical - yet existing systems treat data quality control (QC) as an isolated preprocessin…